返回
Hindsight-aware deep reinforcement learning algorithm for multi-agent systems
DOI:10.1007/s13042-022-01505-x.png)
摘要
En 中文
Classic reinforcement learning algorithms generate experiences by the agent's constant trial and error, which leads to a large number of failure experiences stored in the replay buffer. As a result, the agents can only learn through these low-quality experiences. In the case of multi-agent systems, this problem is more serious. MADDPG (Multi-Agent Deep Deterministic Policy Gradient) has achieved significant results in solving multi-agent problems by using a framework of centralized training with decentralized execution. Nevertheless, the problem of too many failure experiences in the replay buffer has not been resolved. In this paper, we propose HMADDPG (Hindsight Multi-Agent Deep Deterministic Policy Gradient) to mitigate the negative impact of failure experience. HMADDPG has a hindsight unit, which allows the agents to reflect and produces pseudo experiences that tend to succeed. Pseudo experiences are stored in the replay buffer, so that the agents can combine two kinds of experiences to learn. We have evaluated our algorithm on a number of environments. The results show that the algorithm can guide agents to learn better strategies and can be applied in multi-agent systems which are cooperative, competitive, or mixed cooperative and competitive.
Keyword:
Artificial intelligence
Machine learning
Multi-agent system
Hindsight
Reinforcement learning
Experience replay
期刊
IF:
2.7
论文数:
3.2K
被引数:
5.6K
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Guided goal generation for hindsight multi-goal reinforcement learning事后多目标强化学习的引导式目标生成
NEUROCOMPUTING
IF6.5
Path Planning for Multi-Arm Manipulators Using Deep Reinforcement Learning: Soft Actor-Critic with Hindsight Experience Replay使用深度强化学习的多臂操纵器的路径规划: 具有事后经验重播的软行动者-批评家
SENSORS
IF3.5
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis撒哈拉以南非洲非正规工人的信通技术: 系统回顾和分析
Learning Effective Skeletal Representations on RGB Video for Fine-Grained Human Action Quality Assessment
Electronics
IF0
A Distributed Electricity Trading System in Active Distribution Networks Based on Multi-Agent Coalition and Blockchain基于多Agent联盟和区块链的主动配电网分布式电力交易系统

