返回
Multiagent Trust Region Policy Optimization
DOI:10.1109/TNNLS.2023.3265358.png)
摘要
En 中文
We extend trust region policy optimization (TRPO) to cooperative multiagent reinforcement learning (MARL) for partially observable Markov games (POMGs). We show that the policy update rule in TRPO can be equivalently transformed into a distributed consensus optimization for networked agents when the agents' observation is sufficient. By using a local convexification and trust-region method, we propose a fully decentralized MARL algorithm based on a distributed alternating direction method of multipliers (ADMM). During training, agents only share local policy ratios with neighbors via a peer-to-peer communication network. Compared with traditional centralized training methods in MARL, the proposed algorithm does not need a control center to collect global information, such as global state, collective reward, or shared policy and value network parameters. Experiments on two cooperative environments demonstrate the effectiveness of the proposed method.
Keyword:
Optimization
Training
Approximation algorithms
Convergence
Games
Gradient methods
Observability
Decentralized learning
multiagent reinforcement learning (MARL)
partially observable
trust region policy optimization (TRPO)
期刊
IF:
8.9
论文数:
7.5K
被引数:
7.2W
机构
引用论文
Event-Triggered Communication Network With Limited-Bandwidth Constraint for Multi-Agent Reinforcement Learning具有有限带宽约束的事件触发通信网络,用于多智能体强化学习

