返回
Improving proximal policy optimization with alpha divergence
DOI:10.1016/j.neucom.2023.02.008.png)
摘要
En 中文
Proximal policy optimization (PPO) is a recent advancement in reinforcement learning, which is formu-lated as an unconstrained optimization problem including two terms: accumulative discount return and Kullback-Leibler (KL) divergence. Currently, there are three PPO versions: primary, adaptive, and clip-ping. The most widely used PPO algorithm is the clipping version, in which the KL divergence is replaced by a clipping function to measure the difference between two policies indirectly. In this paper, we revisit this primary PPO and improve it in two aspects. One is to reformulate it as a linearly combined form to control the trade-off between two terms. The other is to substitute a parametric alpha divergence for KL divergence to measure the difference of two policies more effectively. This novel PPO variant is referred to as alphaPPO in this paper. Experiments on six benchmark environments verify the effectiveness of our alphaPPO, compared with clipping and combined PPOs. CO 2023 Published by Elsevier B.V.
Keyword:
Reinforcement learning
Deep neural networks
Proximal policy optimization
KL divergence
Alpha divergence
Markov decision process
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Physical Activity in Community‐Dwelling Stroke Survivors and a Healthy Population Is Not Explained by Motor Function Only
PM&R
IF0
An Empirical Investigation of Early Stopping Optimizations in Proximal Policy Optimization
IEEE ACCESS
IF3.6
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究

