1
Return

Improving value function decomposition in cooperative multi-agent reinforcement learning

delete2026-08-06
delete0
PRE
AI
H
Hao Qin
H
Hong Chen *
DOI:10.1016/j.neucom.2026.134711delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Value decomposition methods have achieved strong performance in cooperative multi-agent reinforcement learning (MARL) under the centralized training with decentralized execution paradigm. However, effectively incorporating explicit, sparse, and time-varying coordination structures into monotonic value decomposition remains challenging. Although recent graph-based MARL methods have explored dynamic and group-aware interaction modeling, it is still non-trivial to construct lightweight and interpretable coordination graphs from local observations while preserving stable utility learning and decentralized execution. To address this issue, we propose the Dynamic Graph Attention Mixing Network (DGAT-MIX), a value decomposition framework for explicit dynamic coordination modeling. DGAT-MIX constructs an observation-driven coordination graph from local visibility and distance cues and performs masked multi-head graph attention over the resulting graph. The graph-enhanced representations are then injected into individual utilities through a residual coordination signal before monotonic mixing. This design introduces a sparse and interpretable relational inductive bias while retaining the original utility pathway for stable optimization. Experiments on MPE and SMACv2, together with additional analyses on selected SMAC scenarios, show that DGAT-MIX achieves competitive or superior performance compared with representative baselines, especially in heterogeneous, asymmetric, and dynamically changing cooperative tasks.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

B
beijing union university
Scholars:
260
Papers: 102
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers