arrow
返回

Safe Reinforcement Learning Using Robust MPC

delete2021-08-01
delete155
delete
OA
AI
M
Mario Zanon *
S
Sébastien Gros
DOI:10.1109/TAC.2020.3024161delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Reinforcement learning (RL) has recently impressed the world with stunning results in various applications. While the potential of RL is now well established, many critical aspects still need to be tackled, including safety and stability issues. These issues, while secondary for the RL community, are central to the control community that has been widely investigating them. Model predictive control (MPC) is one of the most successful control techniques because, among others, of its ability to provide such guarantees even for uncertain constrained systems. Since MPC is an optimization-based technique, optimality has also often been claimed. Unfortunately, the performance of MPC is highly dependent on the accuracy of the model used for predictions. In this article, we propose to combine RL and MPC in order to exploit the advantages of both, and therefore, obtain a controller that is optimal and safe. We illustrate the results with two numerical examples in simulations.
Keyword:
Safety
Robustness
Data models
Numerical models
Uncertainty
Stability analysis
Computational modeling
Reinforcement learning (RL)
robust model predictive control (MPC)
safe policies
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

I
IMT School for Advanced Studies Lucca
学者数:
676
论文数: 701
被引数: 693
引用论文

引用论文

Economic optimization using model predictive control with a terminal cost
err2011-12-01
err349
PREAI
errAmrit, Rishi; Rawlings, James B.; Angeli, David
err分享
err收藏
A tracking MPC formulation that is locally equivalent to economic MPC
err2016-09-01
err34
PREAI
errZanon, Mario; Gros, Sebastien; Diehl, Moritz
err分享
err收藏
Ribosome protection by ABC‐F proteins—Molecular mechanism and potential drug design
err2019-03-04
err0
errOAAI
errRya Ero; Veerendra Kumar; Weixin Su; Yong‐Gui Gao
err分享
err收藏
学者 查看更多内容