arrow
Return

Data Efficient Deep Reinforcement Learning With Action-Ranked Temporal Difference Learning

delete2024-08-01
delete1
PRE
AI
刘琪 (Qi Liu)
李彦杰 cover
李彦杰 (Yanjie Li) *
Y
Yuecheng Liu
K
Ke Lin
J
Jianqi Gao
Y
Yunjiang Lou
DOI:10.1109/TETCI.2024.3369641delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In value-based deep reinforcement learning (RL), value function approximation errors lead to suboptimal policies. Temporal difference (TD) learning is one of the most important methodologies to approximate state-action (Q) value function. In TD learning, it is critical to estimate Q values of greedy actions more accurately because a more accurate target Q value enhances the estimation accuracy of Q value. To improve the estimation accuracy of Q value, we propose an action-ranked TD learning method to enhance the performance of deep RL by weighting each TD error according to the rank of its corresponding state-action pair's value among all the Q values on a state. The proposed method can provide more accurate target values for TD learning, making the estimation of the Q value more accurate. We apply the proposed method to a representative value-based deep RL algorithm, and results show that the proposed method outperforms baselines on 31 out of 40 Atari games. Furthermore, we extend the proposed method to multi-agent deep RL. To adaptively determine the hyperparameter in action-ranked TD learning, we propose a meta action-ranked TD learning. A series of experiments quantitatively verify that our methods outperform baselines on Atari games, StarCraft-II, and Grid World environments.
Keywords:
Reinforcement learning
temporal difference
data efficient
action-rank
meta learning

Journal

I
IEEE Transactions on Emerging Topics in Computational Intelligence
IF:
6.5
Papers:
1.4K
Citations:
4.5K

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66