arrow
Return

Alternated Greedy-Step Deterministic Policy Gradient

delete2023-12-01
delete1
PRE
AI
X
Xuesong Wang
J
Jiazhi Zhang
Y
Yang Gu
L
Longyang Huang
K
Kun Yu
Y
Yuhu Cheng *
DOI:10.1109/TCDS.2023.3242274delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The greedy-step $Q$ -learning (GQL) can effectively accelerate the $Q$ -value updating process. However, since it is an improved version of $Q$ -learning, the problem of $Q$ -value overestimation also exists. Since there are in total two max operators used to iteratively calculate $Q$ -value in GQL, many existing solutions to reduce the $Q$ -value estimation bias are invalid for GQL. To address the issue, an alternated greedy-step update (AGU) framework that consists of two independent $Q$ -value estimators is proposed in this study. In the proposed AGU framework, one estimator is to determine the time step that can maximize the estimated $n$ -step return and the other estimator is to update the prior estimator using the target value calculated on the basis of the determined time step. The convergence of the AGU framework is proved in theoretical. In addition, an alternated greedy-step deterministic policy gradient (AGDPG) that can be applied to continuous-action tasks is proposed by combining the AGU framework with deep deterministic policy gradient (DDPG). Experiments of AGDPG on continuous-action tasks of the MuJoCo platform highlight its superior performance.
Keywords:
Alternated greedy-step update (AGU)
deep deterministic policy gradient (DDPG)
greedy-step Q-learning (GQL)
Q-value overestimation

Journal

IEEE Transactions on Cognitive and Developmental Systems cover
IEEE Transactions on Cognitive and Developmental Systems
IF:
4.9
Papers:
1.0K
Citations:
3.5K

Organization

No organization information available