arrow
返回

DVF:Multi-agent Q-learning with difference value factorization

delete2024-02-01
delete0
PRE
AI
A
Anqi Huang
Y
Yongli Wang *
S
Sang, Jianghui
W
Wang, Xiaoli
Y
Yupeng Wang
DOI:10.1016/j.knosys.2024.111422delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In cooperative multi -agent games, agents are required to learn effective cooperative behaviors within complex action spaces. An effective approach is to utilize the Individual -Global -Max (IGM) principle to decompose the global reward signal into individual contributions from each agent. However, existing methods confront two significant challenges: ensuring the accuracy of the value function decomposition process and guaranteeing monotonic improvement in joint policies. These challenges become exacerbated in scenarios involving nonmonotonic matrices. To address these challenges, we introduce a novel and flexible value factorization method called Difference Value Factorization (DVF). The key idea of our method is to transform the IGM principle into a new form DVF-IGM, which addresses non -monotonic constraints by ensuring the consistency of the IGM process between the joint difference value and the complex non-linear sum of independent difference values. A centralized evaluator is employed to estimate global Q -values, which not only enhances expressiveness but also constructs difference values for updating individual value functions. We demonstrate that DVF-IGM is an equivalent transformation of IGM and that DVF has the monotonic improvement property. Empirically, our method has been shown to maintain and recover the optimal policy in non -monotonic matrix games and achieve state-of-the-art performance in cooperative tasks within the StarCraft Multi -Agent Challenge (SMAC).
Keyword:
Multi-agent reinforcement learning
Value factorization
Individual-global-max
Reinforcement learning
Multi-agent system

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

暂无机构信息
引用论文

引用论文

ET-HF: A novel information sharing model to improve multi-agent cooperation
err2022-12-01
err9
PREAI
errXie, Shaorong; Zhang, Han; Yu, Hang; Li, Yang; Zhang, Zhenyu; Luo, Xiangfeng
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
Priming
err2002-01-01
err0
PREAI
errAnthony D. Wagner; Wilma Koutstaal
err分享
err收藏
Changes in muscle T2 and tissue damage after downhill running in mdx Mice
err2011-04-12
err0
errOAAI
errSunita Mathur; Ravneet S. Vohra; Sean A. Germain; Sean Forbes; Nathan D. Bryant; Krista Vandenborne; Glenn A. Walter
err分享
err收藏
学者 查看更多内容