Return
MABQN: Multi-agent reinforcement learning algorithm with discrete policy
DOI:10.1016/j.neucom.2025.129552.png)
Abstract
En 中文
Cooperative multi-agent reinforcement learning (MARL) for continuous control has diverse applications in real-world scenarios. Most of those MARL algorithms focus on enhancing performance through policy-based paradigms, the challenge of low sample efficiency caused by the continuous nature remains underexplored. To address this issue, we propose the Multi-Agent Branching Q-Networks (MABQN) algorithm, an improved QMIX architecture integrating action discretization and value decomposition. MABQN reduces the policy search space by progressively discretizing the continuous action space and decoupling action dimensions, thereby improving learning efficiency. Moreover, it employs a centralized hypernetwork to decompose joint action values, mitigating the credit assignment problem. Experimental results demonstrate that MABQN outperforms other mainstream cooperative MARL algorithms across continuous, discrete, and hybrid action space tasks.
Keywords:
Multi-agent reinforcement learning
Progressive action discretization
Discrete control
Value decomposition

