arrow
Return

Stochastic Double Deep Q-Network

delete2019-01-01
delete20
delete
OA
AI
P
Pingli Lv
X
Xuesong Wang *
Y
Yuhu Cheng
Z
Ziming Duan
DOI:10.1109/ACCESS.2019.2922706delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Estimation bias seriously affects the performance of reinforcement learning algorithms. The maximum operation may result in overestimation, while the double estimator operation often leads to underestimation. To eliminate the estimation bias, these two operations are combined together in our proposed algorithm named stochastic double deep Q-learning network (SDDQN), which is based on the idea of random selection. A tabular version of SDDQN is also given, named stochastic double Q-learning (SDQ). Both the SDDQN and SDQ are based on the double estimator framework. At each step, we choose to use either the maximum operation or the double estimator operation with a certain probability, which is determined by a random selection parameter. The theoretical analysis shows that there indeed exists a proper random selection parameter that makes SDDQN and SDQ unbiased. The experiments on Grid World and Atari 2600 games illustrate that our proposed algorithms can balance the estimation bias effectively and improve performance.
Keywords:
Estimation bias
deep reinforcement learning
maximum operation
double estimator operation
stochastic combination
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

No organization information available
Cited Papers

Cited Papers

Distributed Services Attestation in IoT
err2018-11-30
err0
PREAI
errMauro Conti; Edlira Dushku; Luigi V. Mancini
errShare
errSave
Finite-time analysis of the multiarmed bandit problem
err2002-01-01
err4.2K
PREAI
errAuer, P; Cesa-Bianchi, N; Fischer, P
errShare
errSave
MAP3738c and MptD are specific tags of Mycobacterium avium subsp. paratuberculosis infection in type I diabetes mellitus
err2011-10-01
err0
PREAI
errAndrea Cossu; Valentina Rosu; Daniela Paccagnini; Davide Cossu; Adolfo Pacifico; Leonardo Antonio Sechi
errShare
errSave
Actor-Critic Off-Policy Learning for Optimal Control of Multiple-Model Discrete-Time Systems
err2018-01-01
err48
errOAAI
errSkach, Jan; Kiumarsi, Bahare; Lewis, Frank L.; Straka, Ondrej
errShare
errSave
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis
err2017-09-01
err0
PREAI
errNasibu Mramba; Joel Rumanyika; Mikko Apiola; Jarkko Suhonen
errShare
errSave
researcher View more