arrow
Return

Sampling Efficient Deep Reinforcement Learning Through Preference-Guided Stochastic Exploration

delete2024-12-01
delete4
delete
OA
AI
H
Huang, Wenhui
C
Cong Zhang
J
Jingda Wu
X
Xiangkun He
J
Jie Zhang
C
Chen Lv *
DOI:10.1109/TNNLS.2023.3317628delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Stochastic exploration is the key to the success of the deep Q-network (DQN) algorithm. However, most existing stochastic exploration approaches either explore actions heuristically regardless of their Q values or couple the sampling with Q values, which inevitably introduce bias into the learning process. In this article, we propose a novel preference-guided epsilon-greedy exploration algorithm that can efficiently facilitate exploration for DQN without introducing additional bias. Specifically, we design a dual architecture consisting of two branches, one of which is a copy of DQN, namely, the Q branch. The other branch, which we call the preference branch, learns the action preference that the DQN implicitly follows. We theoretically prove that the policy improvement theorem holds for the preference-guided epsilon-greedy policy and experimentally show that the inferred action preference distribution aligns with the landscape of corresponding Q values. Intuitively, the preference-guided epsilon-greedy exploration motivates the DQN agent to take diverse actions, so that actions with larger Q values can be sampled more frequently, and those with smaller Q values still have a chance to be explored, thus encouraging the exploration. We comprehensively evaluate the proposed method by benchmarking it with well-known DQN variants in nine different environments. Extensive results confirm the superiority of our proposed method in terms of performance and convergence speed.
Keywords:
Data efficiency
deep Q network (DQN)
deep reinforcement learning (DRL)
preference-guided exploration
stochastic policy

Journal

IEEE Transactions on Neural Networks and Learning Systems cover
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
Papers:
7.5K
Citations:
7.2W

Organization

N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W