arrow
返回

Distributed entropy-regularized multi-agent reinforcement learning with policy consensus

delete2024-06-01
delete1
PRE
AI
Y
Yifan Hu
付俊杰 (Junjie Fu) *
G
Guanghui Wen
吕跃祖 (Yuezu Lv)
W
Wei Ren
DOI:10.1016/j.automatica.2024.111652delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Sample efficiency is a limiting factor for existing distributed multi -agent reinforcement learning (MARL) algorithms over networked multi -agent systems. In this paper, the sample efficiency problem is tackled by formally incorporating the entropy regularization into the distributed MARL algorithm design. Firstly, a new entropy -regularized MARL problem is formulated under the model of networked multi -agent Markov decision processes with observation -based policies and homogeneous agents, where the policy parameter sharing among the agents provably preserves the optimality. Secondly, an on -policy distributed actor-critic algorithm is proposed, where each agent shares its parameters of both the critic and actor for consensus update. Then, the convergence analysis of the proposed algorithm is provided based on the stochastic approximation theory under the assumption of linear function approximation of the critic. Furthermore, a practical off -policy version of the proposed algorithm is developed which possesses scalability, data efficiency and learning stability. Finally, the proposed distributed algorithm is compared against the solid baselines including two classic centralized training algorithms in the multi -agent particle environment, whose learning performance is empirically demonstrated through extensive simulation experiments. (c) 2024 Elsevier Ltd. All rights reserved.
Keyword:
Distributed actor-critic algorithm
Networked multi-agent system
Entropy regularization
Deep reinforcement learning

期刊

Automatica 封面图
Automatica
IF:
5.9
论文数:
1.2W
被引数:
5.2W

机构

B
beijing institute of technology
学者数:
5.5W
论文数: 4.0W
被引数: 63
S
southeast university - china
学者数:
5.3W
论文数: 4.9W
被引数: 57
University of California System 封面图
University of California System
学者数:
37.5W
论文数: 33.7W
被引数: 6.6K
学者 查看更多机构
引用论文

引用论文

Graphs, Convolutions, and Neural Networks: From Graph Filters to Graph Neural Networks
err2020-11-01
err112
errOAAI
errGama, Fernando; Isufi, Elvin; Leus, Geert; Ribeiro, Alejandro
err分享
err收藏
err分享
err收藏
Joint Robustness of Time-Varying Networks and Its Applications to Resilient Consensus
err2023-11-01
err13
PREAI
errWen, Guanghui; Lv, Yuezu; Zheng, Wei Xing; Zhou, Jialing; Fu, Junjie
err分享
err收藏
Analysis on smart material suitable for autogenous microelectronic application
err2019-08-30
err0
PREAI
errR Sitharthan; Manimaran Ponnusamy; Madurakavi Karthikeyan; D Shanmuga Sundar
err分享
err收藏
The use of mucoadhesive oral patches containing epigallocatechin-3-gallate to treat periodontitis: an in vivo study
err2022-12-01
err0
errOAAI
errDur Muhammad Lashari; Mohammed Aljunaid; Yasmeen Lashari; Huda Rashad Qaid; Rini Devijanti Ridwan; Indeswati Diyatri; Nejva Kaid; Baleegh Abdulraoof Alkadasi
err分享
err收藏
学者 查看更多内容