arrow
返回

Hindsight-aware deep reinforcement learning algorithm for multi-agent systems

delete2022-01-29
delete3
PRE
AI
C
Chengjing Li
L
Li Wang *
DOI:10.1007/s13042-022-01505-xdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Classic reinforcement learning algorithms generate experiences by the agent's constant trial and error, which leads to a large number of failure experiences stored in the replay buffer. As a result, the agents can only learn through these low-quality experiences. In the case of multi-agent systems, this problem is more serious. MADDPG (Multi-Agent Deep Deterministic Policy Gradient) has achieved significant results in solving multi-agent problems by using a framework of centralized training with decentralized execution. Nevertheless, the problem of too many failure experiences in the replay buffer has not been resolved. In this paper, we propose HMADDPG (Hindsight Multi-Agent Deep Deterministic Policy Gradient) to mitigate the negative impact of failure experience. HMADDPG has a hindsight unit, which allows the agents to reflect and produces pseudo experiences that tend to succeed. Pseudo experiences are stored in the replay buffer, so that the agents can combine two kinds of experiences to learn. We have evaluated our algorithm on a number of environments. The results show that the algorithm can guide agents to learn better strategies and can be applied in multi-agent systems which are cooperative, competitive, or mixed cooperative and competitive.
Keyword:
Artificial intelligence
Machine learning
Multi-agent system
Hindsight
Reinforcement learning
Experience replay

期刊

International Journal of Machine Learning and Cybernetics 封面图
International Journal of Machine Learning and Cybernetics
IF:
2.7
论文数:
3.2K
被引数:
5.6K

机构

T
Taiyuan University of Technology
学者数:
2.2W
论文数: 1.4W
被引数: 1.8W
引用论文

引用论文

Priming
err2002-01-01
err0
PREAI
errAnthony D. Wagner; Wilma Koutstaal
err分享
err收藏
Learning Effective Skeletal Representations on RGB Video for Fine-Grained Human Action Quality Assessment
err2020-03-28
err0
errOAAI
errQing Lei; Hong-Bo Zhang; Ji-Xiang Du; Tsung-Chih Hsiao; Chih-Cheng Chen
err分享
err收藏
err分享
err收藏
Genetic basis of lacunar stroke: a pooled analysis of individual patient data and genome-wide association studies
err2021-05-01
err0
errOAAI
errMatthew Traylor; Elodie Persyn; Liisa Tomppo; Sofia Klasson; Vida Abedi; Mark K Bakker; Nuria Torres; Linxin Li; Steven Bell; Loes Rutten-Jacobs; Daniel J Tozer; Christoph J Griessenauer; Yanfei Zhang; Annie Pedersen; Pankaj Sharma; Jordi Jimenez-Conde; Tatjana Rundek; Raji P Grewal; Arne Lindgren; James F Meschia; Veikko Salomaa; Aki Havulinna; Christina Kourkoulis; Katherine Crawford; Sandro Marini; Braxton D Mitchell; Steven J Kittner; Jonathan Rosand; Martin Dichgans; Christina Jern; Daniel Strbian; Israel Fernandez-Cadenas; Ramin Zand; Ynte Ruigrok; Natalia Rost; Robin Lemmens; Peter M Rothwell; Christopher D Anderson; Joanna Wardlaw; Cathryn M Lewis; Hugh S Markus
err分享
err收藏
学者 查看更多内容