arrow
返回

MP-TD3: Multi-Pool Prioritized Experience Replay-Based Asynchronous Twin Delayed Deep Deterministic Policy Gradient Algorithm

delete2024-01-01
delete1
delete
OA
AI
D
Detian Huang *
DOI:10.1109/ACCESS.2024.3435949delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The prioritized experience replay mechanisms have achieved remarkable success in accelerating the convergence of reinforcement learning algorithms. However, applying traditional prioritized experience replay mechanisms directly to asynchronous reinforcement learning leads to slow convergence, due to the difficulty for an agent to utilize excellent experiences obtained by other agents interacting with the environment. To address the above issue, we propose a Multi-pool Prioritized experience replay-based asynchronous Twin Delayed Deep Deterministic policy gradient algorithm (MP-TD3). Specifically, a multi-pool prioritized experience replay mechanism is proposed to strengthen the experience interactions among different agents to accelerate the network convergence. Then, a global-pool self-cleaning mechanism based on sample diversity and a global-pool self-cleaning mechanism based on TD-errors are designed to overcome the deficiency that the samples suffer from high redundancy and low information content in the global-pool, respectively. Finally, a multi-batch sampling mechanism is investigated to further reduce the training time. Extensive experiments validate that the proposed MP-TD3 significantly improve the convergence speed and performance compared with state-of-the-art methods.
Keyword:
Training
Parallel processing
Convergence
Correlation
Network architecture
Training data
Reinforcement learning
Asynchronous reinforcement learning
twin delayed deep deterministic algorithm
prioritized experience replay
TD-error

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

H
huaqiao university
学者数:
1.1W
论文数: 7.1K
被引数: 131
引用论文

引用论文

Prioritized Experience Replay based on Multi-armed Bandit
err2022-03-01
err10
PREAI
errLiu, Ximing; Zhu, Tianqing; Jiang, Cuiqing; Ye, Dayong; Zhao, Fuqing
err分享
err收藏
TOMATO GROWTH AND YIELD AFFECTED BY NICKEL PRESENTED IN THE NUTRIENT SOLUTION
err1998-04-01
err0
PREAI
errJ. Balaguer; M.B. Almendro; I. Gómez; J. Navarro Pedreño; J. Mataix
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Deep Reinforcement Learning: A Brief Survey深度强化学习: 简要综述
err2017-11-01
err2.4K
errOAAI
errArulkumaran, Kai; Deisenroth, Marc Peter; Brundage, Miles; Bharath, Anil Anthony
err分享
err收藏
err分享
err收藏
学者 查看更多内容