arrow
Return

Regularization-Adapted Anderson Acceleration for multi-agent reinforcement learning

delete2023-09-01
delete3
PRE
AI
S
Siying Wang
W
Wenyu Chen
L
Liwei Huang
F
Fan Zhang
Z
Zhitong Zhao
H
Hong Qu *
DOI:10.1016/j.knosys.2023.110709delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Originating from model-free reinforcement learning (RL), many modern multi-agent reinforcement learning (MARL) algorithms are usually armed with the paradigm of Centralized Training with De-centralized Execution (CTDE) to mitigate the non-stationary problem and make the training process stable. However, these methods still suffer from sample inefficiency and slow training convergence as in the single-agent reinforcement learning setting. Many common methods aiming to tackle these problems utilize the experience buffer and parallel training mechanism to speed up the training process, which would cost more computing resources and may still underuse the sampled experiences. In this paper, we propose Regularization -Adapted Anderson Acceleration (RA3) for model-free, off-policy MARL algorithms. Under the CTDE paradigm, this specific RA3 approach treats the joint action-value function update as a fixed-point iteration task and speeds up the training process with the same amount of sampled experiences as in the baseline algorithms. Furthermore, our RA3 employs an adaptive regularization strategy related to Bellman residuals to stabilize the update process and enhance the training performance. Experimental results demonstrate that the improved learning speed and superior performance of our proposed method are significantly improved on the predator-prey game and the challenging StarCraft II micromanagement benchmark tasks.& COPY; 2023 Elsevier B.V. All rights reserved.
Keywords:
Reinforcement learning
Decentralized partially observable Markov
decision process (Dec-POMDP)
Multi-agent reinforcement learning
Q-learning
Anderson acceleration

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

No organization information available