arrow
返回

Improving proximal policy optimization with alpha divergence

delete2023-05-01
delete4
PRE
AI
H
Haotian Xu
Z
Zheng Yan
J
Junyu Xuan
张
张广泉 (Guangquan Zhang)
J
Jie Lü *
DOI:10.1016/j.neucom.2023.02.008delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Proximal policy optimization (PPO) is a recent advancement in reinforcement learning, which is formu-lated as an unconstrained optimization problem including two terms: accumulative discount return and Kullback-Leibler (KL) divergence. Currently, there are three PPO versions: primary, adaptive, and clip-ping. The most widely used PPO algorithm is the clipping version, in which the KL divergence is replaced by a clipping function to measure the difference between two policies indirectly. In this paper, we revisit this primary PPO and improve it in two aspects. One is to reformulate it as a linearly combined form to control the trade-off between two terms. The other is to substitute a parametric alpha divergence for KL divergence to measure the difference of two policies more effectively. This novel PPO variant is referred to as alphaPPO in this paper. Experiments on six benchmark environments verify the effectiveness of our alphaPPO, compared with clipping and combined PPOs. CO 2023 Published by Elsevier B.V.
Keyword:
Reinforcement learning
Deep neural networks
Proximal policy optimization
KL divergence
Alpha divergence
Markov decision process

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

U
university of technology sydney
学者数:
1.6W
论文数: 2.0W
被引数: 25
引用论文

引用论文

Physical Activity in Community‐Dwelling Stroke Survivors and a Healthy Population Is Not Explained by Motor Function Only
err2013-08-23
err0
PREAI
errAnna Danielsson; Cristiane Meirelles; Carin Willen; Katharina Stibrant Sunnerhagen
err分享
err收藏
An Empirical Investigation of Early Stopping Optimizations in Proximal Policy Optimization
err2021-01-01
err16
errOAAI
errDossa, Rousslan Fernand Julien; Huang, Shengyi; Ontanon, Santiago; Matsubara, Takashi
err分享
err收藏
Direct digital design of a sliding mode‐based control of a PWM synchronous buck converter
err2017-10-01
err0
PREAI
errEnric Vidal‐Idiarte; Adria Marcos‐Pastor; Roberto Giral; Javier Calvente; Luis Martinez‐Salamero
err分享
err收藏
err分享
err收藏
Deep Reinforcement Learning for Autonomous Driving: A Survey用于自动驾驶的深度强化学习: 一项调查
err2022-06-01
err965
errOAAI
errKiran, B. Ravi; Sobh, Ibrahim; Talpaert, Victor; Mannion, Patrick; Al Sallab, Ahmad A.; Yogamani, Senthil; Perez, Patrick
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容