arrow
返回

A monotonic policy optimization algorithm for high-dimensional continuous control problem in 3D MuJoCo

delete2018-06-04
delete2
PRE
AI
Q
Qunyong Yuan *
N
Nanfeng Xiao
DOI:10.1007/s11042-018-6098-ydelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
One challenge in applying reinforcement learning with nonlinear function approximator to high- dimensional continuous control problems is that the update policy produced by the many existed algorithms may fail to improve policy performance or even causes a serious degradation of the policy performance. To address this challenge, this paper proposes a new lower bound on the policy improvement where an average policy divergence on state space is penalized. To the best of our knowledge, this is currently the best result about the lower bound on the policy improvement. Optimizing directly the lower bound on the policy improvement is very difficult, because it demands for high computational overhead. According to the ideal of the trust region policy optimization (TRPO), this paper also presents a monotonic policy optimization algorithm, which is based on the new lower bound on the policy improvement introduced in this paper, it can generate a sequence of monotonically improving policies, and it is suitable for the large-scale continuous control problems. This paper also evaluates and compares the proposed algorithms with some of the existed algorithms on highly challenging robot locomotion tasks.
Keyword:
Reinforcement learning
Policy optimization
Continuous control policy
Deep neural network
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
1.9W
被引数:
3.2W

机构

S
south china university of technology
学者数:
6.8W
论文数: 5.1W
被引数: 85