Return
Behavior Cloning-TD3-Based End-to-End Autonomous Flight of Quadrotor UAVs
Y
B
刘
Y
Q
DOI:10.1109/taes.2026.3712715.png)
Abstract
En 中文
This study investigates an end-to-end autonomous flight approach for quadrotor unmanned aerial vehicles under local environmental perception. To address the limitations of traditional hierarchical architectures, such as the decoupling of planning and control, heavy reliance on global perception, an end-to-end flight control method is proposed by integrating behavior cloning (BC) with twin delayed deep deterministic policy gradient. First, onboard light detection and ranging (LiDAR) data are used to perceive the environment and construct a collision risk model, which is combined with positional data to form the state representation. Then, two independent experience storage mechanisms are designed: a prioritized experience replay buffer and a successful trajectory experience pool. After each update of the policy network, the BC model, trained on successful trajectories, is employed to further refine the policy. In addition, a composite reward function with four components dynamic distance reward, heading alignment reward, collision penalty, and time-efficiency penalty is designed to improve the stability and safety of the learned autonomous flight policy. Finally, simulation experiments are conducted to validate the effectiveness and superiority of the proposed approach.
Keywords:
Autonomous flight
behavior cloning (BC) and twin delayed deep deterministic policy gradient (TD3)
end-to-end
local perception
quadrotor unmanned aerial vehicle (UAV)
Journal
IF:
5.7
Papers:
651
Citations:
2.4W
