返回
Data-Based Optimal Consensus Control for Multiagent Systems With Policy Gradient Reinforcement Learning
DOI:10.1109/TNNLS.2021.3054685.png)
摘要
En 中文
This article investigates the optimally distributed consensus control problem for discrete-time multiagent systems with completely unknown dynamics and computational ability differences. The problem can be viewed as solving nonzero-sum games with distributed reinforcement learning (RL), and each agent is a player in these games. First, to guarantee the real-time performance of learning algorithms, a data-based distributed control algorithm is proposed for multiagent systems using offline system interaction data sets. By utilizing the interactive data produced during the run of a real-time system, the proposed algorithm improves system performance based on distributed policy gradient RL. The convergence and stability are guaranteed based on functional analysis and the Lyapunov method. Second, to address asynchronous learning caused by computational ability differences in multiagent systems, the proposed algorithm is extended to an asynchronous version in which executing policy improvement or not of each agent is independent of its neighbors. Furthermore, an actor-critic structure, which contains two neural networks, is developed to implement the proposed algorithm in synchronous and asynchronous cases. Based on the method of weighted residuals, the convergence and optimality of the neural networks are guaranteed by proving the approximation errors converge to zero. Finally, simulations are conducted to show the effectiveness of the proposed algorithm.
Keyword:
Multi-agent systems
Consensus control
Games
Heuristic algorithms
Dynamic programming
Synchronization
Reinforcement learning
Asynchronous learning
data-based control
nonzero-sum games
optimal distributed consensus control
policy gradient (PG) reinforcement learning (RL)
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
Robust cooperative output regulation of multi-agent systems via adaptive event-triggered control基于自适应事件触发控制的多智能体系统鲁棒协同输出调节
AUTOMATICA
IF5.9
Off-Policy Actor-Critic Structure for Optimal Control of Unknown Systems With Disturbances具有扰动的未知系统最优控制的非策略参与者-批评者结构
Optimal Control for Unknown Discrete-Time Nonlinear Markov Jump Systems Using Adaptive Dynamic Programming基于自适应动态规划的未知离散非线性Markov跳变系统最优控制

