Return
Upper confident bound advantage function proximal policy optimization
DOI:10.1007/s10586-022-03742-9.png)
Abstract
En 中文
Proximal Policy Optimization (PPO) is one of the classical and excellent algorithms in Deep Reinforcement Learning (DRL). However, there are still two problems with PPO. The one problem is that PPO limits the policy update to a certain range, which makes PPO prone to the risk of insufficient exploration, the other problem is that PPO adopts mini-batch update method which leads to negative advantage estimation interference. To address these issues, we propose a new model-free algorithm, called Upper Confident Bound Advantage Function Proximal Policy Optimization (UCB-AF), which estimates the confidence of the advantage estimation through Hoeffding's inequality, increases and adjusts advantage estimation with an upper confidence bound. Moreover, compare to PPO in multiple complex environments, our method not only improves the exploration ability, but enjoys better performance bound as well.
Keywords:
Proximal policy optimization (PPO)
Deep reinforcement learning (DRL)
Upper confident bound (UCB)
Advantage function
Hoeffding inequality
Exploration ability
Journal
C
IF:
4.1
Papers:
5.1K
Citations:
7.5K

