返回
Upper confident bound advantage function proximal policy optimization
DOI:10.1007/s10586-022-03742-9.png)
摘要
En 中文
Proximal Policy Optimization (PPO) is one of the classical and excellent algorithms in Deep Reinforcement Learning (DRL). However, there are still two problems with PPO. The one problem is that PPO limits the policy update to a certain range, which makes PPO prone to the risk of insufficient exploration, the other problem is that PPO adopts mini-batch update method which leads to negative advantage estimation interference. To address these issues, we propose a new model-free algorithm, called Upper Confident Bound Advantage Function Proximal Policy Optimization (UCB-AF), which estimates the confidence of the advantage estimation through Hoeffding's inequality, increases and adjusts advantage estimation with an upper confidence bound. Moreover, compare to PPO in multiple complex environments, our method not only improves the exploration ability, but enjoys better performance bound as well.
Keyword:
Proximal policy optimization (PPO)
Deep reinforcement learning (DRL)
Upper confident bound (UCB)
Advantage function
Hoeffding inequality
Exploration ability
期刊
C
IF:
4.1
论文数:
5.0K
被引数:
7.5K
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Chirality sensing of various biomolecules with helical poly(phenylacetylene)s bearing acidic functional groups in water带有酸性官能团的螺旋聚 (苯乙炔) 在水中对各种生物分子的手性传感
Analysis of increased urinary acid glycosaminoglycans in a patient with relapsing polychondritis复发性多发性软骨炎患者尿酸性糖胺聚糖增加的分析
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis撒哈拉以南非洲非正规工人的信通技术: 系统回顾和分析

