arrow
返回

Regularization-Adapted Anderson Acceleration for multi-agent reinforcement learning

delete2023-09-01
delete3
PRE
AI
S
Siying Wang
W
Wenyu Chen
L
Liwei Huang
F
Fan Zhang
Z
Zhitong Zhao
H
Hong Qu *
DOI:10.1016/j.knosys.2023.110709delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Originating from model-free reinforcement learning (RL), many modern multi-agent reinforcement learning (MARL) algorithms are usually armed with the paradigm of Centralized Training with De-centralized Execution (CTDE) to mitigate the non-stationary problem and make the training process stable. However, these methods still suffer from sample inefficiency and slow training convergence as in the single-agent reinforcement learning setting. Many common methods aiming to tackle these problems utilize the experience buffer and parallel training mechanism to speed up the training process, which would cost more computing resources and may still underuse the sampled experiences. In this paper, we propose Regularization -Adapted Anderson Acceleration (RA3) for model-free, off-policy MARL algorithms. Under the CTDE paradigm, this specific RA3 approach treats the joint action-value function update as a fixed-point iteration task and speeds up the training process with the same amount of sampled experiences as in the baseline algorithms. Furthermore, our RA3 employs an adaptive regularization strategy related to Bellman residuals to stabilize the update process and enhance the training performance. Experimental results demonstrate that the improved learning speed and superior performance of our proposed method are significantly improved on the predator-prey game and the challenging StarCraft II micromanagement benchmark tasks.& COPY; 2023 Elsevier B.V. All rights reserved.
Keyword:
Reinforcement learning
Decentralized partially observable Markov
decision process (Dec-POMDP)
Multi-agent reinforcement learning
Q-learning
Anderson acceleration

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

暂无机构信息
引用论文

引用论文

Nonsyndromic Cleft Lip With or Without Cleft Palate in China: Assessment of Candidate Regions
err2002-03-01
err0
PREAI
errMary L. Marazita; L. Leigh Field; Margaret E. Cooper; Rose Tobias; Brion S. Maher; Supakit Peanchitlertkajorn; You-e Liu
err分享
err收藏
Influence of SSTT, Ageing Regime and Stretching on IGC, Complex of Properties and Precipitation Behavior of 6013 Alloy
err2000-05-09
err0
PREAI
errV.G. Davydov; V.S. Siniavski; L.B. Ber; K.H. Rendigs; Gerhard Tempus; V.D. Valkov; V.D. Kalinin; Ye.V. Titkova; O.G. Ukolova; Ye.A. Lukina; Ye.I. Shvechkov; Ye.Ya. Kaputkin
err分享
err收藏
Priming
err2002-01-01
err0
PREAI
errAnthony D. Wagner; Wilma Koutstaal
err分享
err收藏
学者 查看更多内容