arrow
Return

Upper confident bound advantage function proximal policy optimization

delete2022-09-14
delete2
PRE
AI
G
Guiliang Xie
W
Wei Zhang *
Z
Zhi Hu
G
Gaojian Li
DOI:10.1007/s10586-022-03742-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Proximal Policy Optimization (PPO) is one of the classical and excellent algorithms in Deep Reinforcement Learning (DRL). However, there are still two problems with PPO. The one problem is that PPO limits the policy update to a certain range, which makes PPO prone to the risk of insufficient exploration, the other problem is that PPO adopts mini-batch update method which leads to negative advantage estimation interference. To address these issues, we propose a new model-free algorithm, called Upper Confident Bound Advantage Function Proximal Policy Optimization (UCB-AF), which estimates the confidence of the advantage estimation through Hoeffding's inequality, increases and adjusts advantage estimation with an upper confidence bound. Moreover, compare to PPO in multiple complex environments, our method not only improves the exploration ability, but enjoys better performance bound as well.
Keywords:
Proximal policy optimization (PPO)
Deep reinforcement learning (DRL)
Upper confident bound (UCB)
Advantage function
Hoeffding inequality
Exploration ability

Journal

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
Papers:
5.1K
Citations:
7.5K

Organization

S
Shanghai University of Engineering Science
Scholars:
7.9K
Papers: 4.9K
Citations: 6.0K
Cited Papers

Cited Papers

Chirality sensing of various biomolecules with helical poly(phenylacetylene)s bearing acidic functional groups in water
err2006-07-21
err0
PREAI
errHisanari Onouchi; Takashi Hasegawa; Daisuke Kashiwagi; Hiroyuki Ishiguro; Katsuhiro Maeda; Eiji Yashima
errShare
errSave
Effect of different nanofillers on non-isothermal crystallization kinetics and electric conductivity of dynamically-vulcanized PP-EPDM Blends
err2016-01-05
err0
PREAI
errBin Yang; Lei Hu; Ru Xia; Fang Chen; Shu-Chun Zhao; Yan-Li Deng; Ming Cao; Jia-Sheng Qian; Peng Chen
errShare
errSave
Safety of Early Pharmacological Thromboprophylaxis after Subarachnoid Hemorrhage
err2014-10-30
err0
errOAAI
errAirton Leonardo de Oliveira Manoel; David Turkel-Parrella; Menno Germans; Ekaterina Kouzmina; Priscila da Silva Almendra; Thomas Marotta; Julian Spears; Simon Abrahamson
errShare
errSave
Residual Sarsa algorithm with function approximation
err2017-11-10
err1
PREAI
errFu Qiming; Hu Wen; Liu Quan; Luo Heng; Hu Lingyao; Chen Jianping
errShare
errSave
errShare
errSave
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis
err2017-09-01
err0
PREAI
errNasibu Mramba; Joel Rumanyika; Mikko Apiola; Jarkko Suhonen
errShare
errSave
researcher View more