返回
ME-MADDPG: An efficient learning-based motion planning method for multiple agents in complex environments
DOI:10.1002/int.22778.png)
摘要
En 中文
Developing efficient motion policies for multiagents is a challenge in a decentralized dynamic situation, where each agent plans its own paths without knowing the policies of the other agents involved. This paper presents an efficient learning-based motion planning method for multiagent systems. It adopts the framework of multiagent deep deterministic policy gradient (MADDPG) to directly map partially observed information to motion commands for multiple agents. To improve the efficiency of MADDPG in sample utilization, so as to train more brilliant agents that can adapt to more complex environments, a strategy named mixed experience (ME) is introduced to MADDPG, and this has led to our proposed ME-MADDPG algorithm. The novel ME strategy can be embodied into three specific mechanisms: (1) an artificial potential field-based sample generator to produce high-quality samples in the early training stage; (2) a dynamic mixed sampling strategy to mix the training data from different sources with a variable proportion; (3) a delayed learning skill to stabilize the training of the multiple agents. A series of experiments have been conducted to verify the performance of the proposed ME-MADDPG algorithm, and it has been demonstrated that, compared with MADDPG, the proposed algorithm can significantly improve the convergence speed and convergence effect in the training process, and it has also shown better efficiency and better adaptability in complex dynamic environments while it is used for multiagent motion planning applications.
Keyword:
deep reinforcement learning
MADDPG
motion planning
multiagent
期刊
IF:
3.7
论文数:
3.1K
被引数:
8.1K
机构
引用论文
The Efficacy of Transurethral Resection of the Prostate in the Patients with Weak Bladder Contractility Index
Urology
IF0
Using live stream technology to conduct workplace observation assessment of trainee dental nurses: an evaluation of effectiveness and user experience
BDJ Open
IF0
Lipid-Binding Activity of Intrinsically Unstructured Cytoplasmic Domains of Multichain Immune Recognition Receptor Signaling Subunits
Biochemistry
IF0
Joint Optimization of Multi-UAV Target Assignment and Path Planning Based on Multi-Agent Reinforcement Learning基于多Agent强化学习的多无人机目标分配与航迹规划联合优化
IEEE ACCESS
IF3.6

