返回
Novel task decomposed multi-agent twin delayed deep deterministic policy gradient algorithm for multi-UAV autonomous path planning
DOI:10.1016/j.knosys.2024.111462.png)
摘要
En 中文
Path planning is one of the most essential parts of task planning. However, multiple unmanned aerial vehicles (UAVs) path planning is a challenge when considering the cooperativity of multiple UAVs and the uncertainty of environments. This study proposed the novel task decomposed multi-agent twin delayed deep deterministic policy gradient (TD-MATD3) algorithm that enables UAVs execute path planning in complex multiple obstacles environments. TD-MATD3 improves upon the multi-agent twin delayed deep deterministic policy gradient (MATD3) algorithm by decomposing path planning task into the navigation task module for flying to the target and the obstacle avoidance task module for avoiding obstacles and other UAVs. Specifically, TD-MATD3 decomposes the Actor-Critic network structure of MATD3 into two corresponding parts according to the reward functions of two task modules. And the navigation features output by the Actor-Critic network of the navigation task module are input to the Actor-Critic network of the obstacle avoidance task module to guide UAVs to complete the overall path planning task. A novel reward function is also proposed to facilitate convergence of the algorithm. Experimental results indicate that TD-MATD3 can effectively accelerate convergence and enhance convergence effect during the training process, and it achieves a higher success rate in complex dynamic environments than multi-agent deep deterministic policy gradient (MADDPG) and MATD3 for multi-UAV path planning problem.
Keyword:
Multiple UAVs
DRL
Path planning
MATD3
Decomposed Actor -Critic network
期刊
K
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Cooperative path planning optimization for multiple UAVs with communication constraints通信约束下多无人机协同航迹规划优化
Multiobjective UAV Path Planning for Emergency Information Collection and Transmission面向应急信息采集与传输的多目标无人机路径规划
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
ME-MADDPG: An efficient learning-based motion planning method for multiple agents in complex environmentsMe-maddpg: 复杂环境中多智能体的高效基于学习的运动规划方法
Autonomous navigation of UAV in multi-obstacle environments based on a Deep Reinforcement Learning approach基于深度强化学习的多障碍物环境下无人机自主导航

