arrow
返回

Multiagent Reinforcement Learning With Unshared Value Functions

delete2015-04-01
delete46
PRE
AI
Y
Yujing Hu *
高扬 封面图
高扬 (Yang Gao)
B
Bo An
DOI:10.1109/TCYB.2014.2332042delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
One important approach of multiagent reinforcement learning (MARL) is equilibrium-based MARL, which is a combination of reinforcement learning and game theory. Most existing algorithms involve computationally expensive calculation of mixed strategy equilibria and require agents to replicate the other agents' value functions for equilibrium computing in each state. This is unrealistic since agents may not be willing to share such information due to privacy or safety concerns. This paper aims to develop novel and efficient MARL algorithms without the need for agents to share value functions. First, we adopt pure strategy equilibrium solution concepts instead of mixed strategy equilibria given that a mixed strategy equilibrium is often computationally expensive. In this paper, three types of pure strategy profiles are utilized as equilibrium solution concepts: pure strategy Nash equilibrium, equilibrium-dominating strategy profile, and nonstrict equilibrium-dominating strategy profile. The latter two solution concepts are strategy profiles from which agents can gain higher payoffs than one or more pure strategy Nash equilibria. Theoretical analysis shows that these strategy profiles are symmetric meta equilibria. Second, we propose a multi-step negotiation process for finding pure strategy equilibria since value functions are not shared among agents. By putting these together, we propose a novel MARL algorithm called negotiation-based Q-learning (NegoQ). Experiments are first conducted in grid-world games, which are widely used to evaluate MARL algorithms. In these games, NegoQ learns equilibrium policies and runs significantly faster than existing MARL algorithms (correlated Q-learning and Nash Q-learning). Surprisingly, we find that NegoQ also performs well in team Markov games such as pursuit games, as compared with team-task-oriented MARL algorithms (such as friend Q-learning and distributed Q-learning).
Keyword:
Game theory
multiagent reinforcement learning
Nash equilibrium
negotiation
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

N
Nanyang Technological University
学者数:
4.9W
论文数: 4.8W
被引数: 8.1W
N
nanjing university
学者数:
7.8W
论文数: 5.6W
被引数: 87
引用论文

引用论文

Specificity of the Carboxypeptidase Inhibitor from Potatoes
err1981-04-01
err0
errOAAI
errG. Michael Hass; Scott P. Ager; Duane Le Tourneau; Judith E. Derr-Makus; Donald J. Makus
err分享
err收藏
Mini‐Mental State Examination
err2002-04-30
err0
PREAI
errJoseph R. Cockrell; Marshal F. Folstein
err分享
err收藏
err分享
err收藏
Impact of Social Punishment on Cooperative Behavior in Complex Networks
err2013-10-28
err168
errOAAI
errWang, Zhen; Xia, Cheng-Yi; Meloni, Sandro; Zhou, Chang-Song; Moreno, Yamir
err分享
err收藏
Game Dynamics and Cost of Learning in Heterogeneous 4G Networks
err2012-01-01
err115
PREAI
errKhan, Manzoor Ahmed; Tembine, Hamidou; Vasilakos, Athanasios V.
err分享
err收藏
err分享
err收藏
学者 查看更多内容