arrow
返回

Multiagent Trust Region Policy Optimization

delete2024-09-01
delete6
delete
OA
AI
H
Hepeng Li
H
Haibo He *
DOI:10.1109/TNNLS.2023.3265358delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We extend trust region policy optimization (TRPO) to cooperative multiagent reinforcement learning (MARL) for partially observable Markov games (POMGs). We show that the policy update rule in TRPO can be equivalently transformed into a distributed consensus optimization for networked agents when the agents' observation is sufficient. By using a local convexification and trust-region method, we propose a fully decentralized MARL algorithm based on a distributed alternating direction method of multipliers (ADMM). During training, agents only share local policy ratios with neighbors via a peer-to-peer communication network. Compared with traditional centralized training methods in MARL, the proposed algorithm does not need a control center to collect global information, such as global state, collective reward, or shared policy and value network parameters. Experiments on two cooperative environments demonstrate the effectiveness of the proposed method.
Keyword:
Optimization
Training
Approximation algorithms
Convergence
Games
Gradient methods
Observability
Decentralized learning
multiagent reinforcement learning (MARL)
partially observable
trust region policy optimization (TRPO)

期刊

IEEE Transactions on Neural Networks and Learning Systems 封面图
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
论文数:
7.5K
被引数:
7.2W

机构

U
University of Rhode Island
学者数:
5.0K
论文数: 4.5K
被引数: 6.3K
引用论文

引用论文

Valproic Acid-Induced Acute Pancreatitis and Multiorgan Failure in a Child
err2013-05-01
err0
PREAI
errAyhan Yaman; Tanl Kendirli; Çağlar Ödek; Ömer Bektaş; Zarife Kuloğlu; Meltem Koloğlu; Erdal İnce; Gülhis Deda
err分享
err收藏
Distributed Policy Evaluation Under Multiple Behavior Strategies
err2015-05-01
err69
errOAAI
errValcarcel Macua, Sergio; Chen, Jianshu; Zazo, Santiago; Sayed, Ali H.
err分享
err收藏
Optimal Control Strategies in Delayed Sharing Information Structures
err2011-07-01
err97
errOAAI
errNayyar, Ashutosh; Mahajan, Aditya; Teneketzis, Demosthenis
err分享
err收藏
Distributed Reinforcement Learning via Gossip
err2017-03-01
err42
errOAAI
errMathkar, Adwaitvedant; Borkar, Vivek S.
err分享
err收藏
Towards a wireless and fully-implantable ECoG system
err2013-06-01
err0
PREAI
errE. Tolstosheeva; J. Hoeffmann; J. Pistor; D. Rotermund; T. Schellenberg; D. Boll; T. Hertzberg; V. Gordillo-Gonzalez; S. Mandon; D. Peters-Drolshagen; M. Schneider; K. Pawelzik; A. Kreiter; S. Paul; W. Lang
err分享
err收藏
Explicit and implicit reinforcement learning across the psychosis spectrum.
err2017-07-01
err0
errOAAI
errDeanna M. Barch; Cameron S. Carter; James M. Gold; Sheri L. Johnson; Ann M. Kring; Angus W. MacDonald; Diego A. Pizzagalli; J. Daniel Ragland; Steven M. Silverstein; Milton E. Strauss
err分享
err收藏
Resistance and functional training reduces knee extensor position fluctuations in functionally limited older adults
err2005-09-29
err0
PREAI
errTodd M. Manini; Brian C. Clark; Brian L. Tracy; Jeanmarie Burke; Lori Ploutz-Snyder
err分享
err收藏
err2002-01-01
err0
PREAI
errGlynis Laws; Deborah Gunn
err分享
err收藏
学者 查看更多内容