arrow
Return

MP-TD3: Multi-Pool Prioritized Experience Replay-Based Asynchronous Twin Delayed Deep Deterministic Policy Gradient Algorithm

delete2024-01-01
delete1
delete
OA
AI
D
Detian Huang *
DOI:10.1109/ACCESS.2024.3435949delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The prioritized experience replay mechanisms have achieved remarkable success in accelerating the convergence of reinforcement learning algorithms. However, applying traditional prioritized experience replay mechanisms directly to asynchronous reinforcement learning leads to slow convergence, due to the difficulty for an agent to utilize excellent experiences obtained by other agents interacting with the environment. To address the above issue, we propose a Multi-pool Prioritized experience replay-based asynchronous Twin Delayed Deep Deterministic policy gradient algorithm (MP-TD3). Specifically, a multi-pool prioritized experience replay mechanism is proposed to strengthen the experience interactions among different agents to accelerate the network convergence. Then, a global-pool self-cleaning mechanism based on sample diversity and a global-pool self-cleaning mechanism based on TD-errors are designed to overcome the deficiency that the samples suffer from high redundancy and low information content in the global-pool, respectively. Finally, a multi-batch sampling mechanism is investigated to further reduce the training time. Extensive experiments validate that the proposed MP-TD3 significantly improve the convergence speed and performance compared with state-of-the-art methods.
Keywords:
Training
Parallel processing
Convergence
Correlation
Network architecture
Training data
Reinforcement learning
Asynchronous reinforcement learning
twin delayed deep deterministic algorithm
prioritized experience replay
TD-error

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

H
huaqiao university
Scholars:
1.1W
Papers: 7.1K
Citations: 131
Cited Papers

Cited Papers

Prioritized Experience Replay based on Multi-armed Bandit
err2022-03-01
err10
PREAI
errLiu, Ximing; Zhu, Tianqing; Jiang, Cuiqing; Ye, Dayong; Zhao, Fuqing
errShare
errSave
TOMATO GROWTH AND YIELD AFFECTED BY NICKEL PRESENTED IN THE NUTRIENT SOLUTION
err1998-04-01
err0
PREAI
errJ. Balaguer; M.B. Almendro; I. Gómez; J. Navarro Pedreño; J. Mataix
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
Deep Reinforcement Learning: A Brief Survey
err2017-11-01
err2.4K
errOAAI
errArulkumaran, Kai; Deisenroth, Marc Peter; Brundage, Miles; Bharath, Anil Anthony
errShare
errSave
errShare
errSave
researcher View more