arrow
Return

Efficient Reward Shaping for Multiagent Systems

delete2024-01-01
delete2
PRE
AI
V
Vrushabh S. Donge *
B
Bosen Lian
F
Frank L. Lewis
A
Ali Davoudi
DOI:10.1109/TCNS.2024.3401000delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this article, we address the reward-shaping problem of large-scale multiagent systems (MASs) using inverse reinforcement learning (IRL). The learning MAS does not have prior knowledge of the cost function of the target MAS and aims to reconstruct it based on the target's demonstrations. We propose a scalable model-free IRL algorithm for a large-scale MAS, where dynamic mode decomposition (DMD) extracts dynamic modes and builds a projection matrix. This significantly reduces the data required while retaining the system's essential dynamic information. The proofs of the algorithm's convergence, stability, and nonuniqueness of the state reward weight are presented. The efficacy of our method is validated with a large-scale consensus network, by comparing the required data sizes and computational time for reward shaping with and without DMD.
Keywords:
Artificial neural networks
Heuristic algorithms
Optimal control
Control systems
Network systems
Dimensionality reduction
Stability criteria
Data-driven control
dynamic mode decomposition (DMD)
inverse reinforcement learning (IRL)
large-scale system
optimal control

Journal

IEEE Transactions on Control of Network Systems cover
IEEE Transactions on Control of Network Systems
IF:
5
Papers:
1.6K
Citations:
5.8K

Organization

U
university of texas system
Scholars:
18.5W
Papers: 15.6W
Citations: 210