返回
Strategic Interaction Multi-Agent Deep Reinforcement Learning
DOI:10.1109/ACCESS.2020.3005734.png)
摘要
En 中文
Despite the proliferation of multi-agent deep reinforcement learning (MADRL), most existing typical methods do not scale well to the dynamics of agent populations. And as the population increases, the dimensional explosion of joint state-action and the complex interaction between agents make learning extremely cumbersome, which poses the scalability challenge for MADRL. This paper focuses on the scalability issue of MADRL with homogeneous agents. In a natural population, local interaction is a more feasible mode of interplay rather than global interaction. And inspired by the strategic interaction model in economics, we decompose the value function of each agent into the sum of the expected cumulative rewards of the interaction between the agent and each neighbor. This novel value function is decentralized and decomposable, which enables it to scale well to the dynamic changes in the number of large-scale agents. Hereby, the corresponding strategic interaction reinforcement learning algorithm (SIQ), is proposed to learn the optimal policy of each agent, wherein a neural network is employed to estimate the expected cumulative reward for the interaction between the agent and one of its neighbors. We test the validity of the proposed method in a mixed cooperative-competitive confrontation game through numerical experiments. Furthermore, the scalability comparison experiments illustrate that the scalability of the SIQ algorithm outperforms the independent learning and mean field reinforcement learning algorithms in multiple scenarios with different and dynamically changing numbers.
Keyword:
Multi-agent deep reinforcement learning
scalability
local interaction
large scale
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Lipid-Binding Activity of Intrinsically Unstructured Cytoplasmic Domains of Multichain Immune Recognition Receptor Signaling Subunits
Biochemistry
IF0
AWESOME: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
MACHINE LEARNING
IF2.9
Porosity and liquation cracking of dissimilar Nd:YAG laser welding of SUS304 stainless steel to T2 copperSUS304不锈钢与T2铜的异种Nd:YAG激光焊接的气孔和液化裂纹
Incorporating temperature-leakage interdependency into dynamic voltage scaling for real-time systems
Learning Effective Skeletal Representations on RGB Video for Fine-Grained Human Action Quality Assessment
Electronics
IF0
Unusual vascular focal high-grade arterial stenoses in a young woman with systemic lupus erythematosus and secondary antiphospholipid syndrome
Lupus
IF0

