arrow
返回

Upper confident bound advantage function proximal policy optimization

delete2022-09-14
delete2
PRE
AI
G
Guiliang Xie
W
Wei Zhang *
Z
Zhi Hu
G
Gaojian Li
DOI:10.1007/s10586-022-03742-9delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Proximal Policy Optimization (PPO) is one of the classical and excellent algorithms in Deep Reinforcement Learning (DRL). However, there are still two problems with PPO. The one problem is that PPO limits the policy update to a certain range, which makes PPO prone to the risk of insufficient exploration, the other problem is that PPO adopts mini-batch update method which leads to negative advantage estimation interference. To address these issues, we propose a new model-free algorithm, called Upper Confident Bound Advantage Function Proximal Policy Optimization (UCB-AF), which estimates the confidence of the advantage estimation through Hoeffding's inequality, increases and adjusts advantage estimation with an upper confidence bound. Moreover, compare to PPO in multiple complex environments, our method not only improves the exploration ability, but enjoys better performance bound as well.
Keyword:
Proximal policy optimization (PPO)
Deep reinforcement learning (DRL)
Upper confident bound (UCB)
Advantage function
Hoeffding inequality
Exploration ability

期刊

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
论文数:
5.0K
被引数:
7.5K

机构

S
Shanghai University of Engineering Science
学者数:
7.9K
论文数: 4.9K
被引数: 6.0K
引用论文

引用论文

Effect of different nanofillers on non-isothermal crystallization kinetics and electric conductivity of dynamically-vulcanized PP-EPDM Blends
err2016-01-05
err0
PREAI
errBin Yang; Lei Hu; Ru Xia; Fang Chen; Shu-Chun Zhao; Yan-Li Deng; Ming Cao; Jia-Sheng Qian; Peng Chen
err分享
err收藏
Safety of Early Pharmacological Thromboprophylaxis after Subarachnoid Hemorrhage
err2014-10-30
err0
errOAAI
errAirton Leonardo de Oliveira Manoel; David Turkel-Parrella; Menno Germans; Ekaterina Kouzmina; Priscila da Silva Almendra; Thomas Marotta; Julian Spears; Simon Abrahamson
err分享
err收藏
Residual Sarsa algorithm with function approximation
err2017-11-10
err1
PREAI
errFu Qiming; Hu Wen; Liu Quan; Luo Heng; Hu Lingyao; Chen Jianping
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容