arrow
Return

LLM-PD: A Large Language Model-Driven Policy Distillation-Based Method for Multi-UAV Path Planning

delete2026-01-21
delete0
PRE
AI
L
Lun Tang
J
Jiaming He
X
Xin Liao
Q
Qinghai Liu
Q
Qianbin Chen
DOI:10.1109/JIOT.2026.3656708delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Industrial Internet of Things (IIoT) has accelerated the adoption of multi-uncrewed aerial vehicle (UAV) systems in applications, such as urban inspection and emergency response. However, the effective path planning in dynamic environments remains challenging due to slow convergence, weak coordination, trajectory oscillation, and limited generalization. To address these issues, this article proposes a large language model-driven policy distillation (LLM-PD)-based method for multi-UAV path planning. By integrating LLM-based semantic reasoning with multiagent reinforcement learning (MARL), LLM-PD introduces a three-stage training paradigm—policy distillation, imitation pretraining, and reinforcement-based self-optimization—to enhance the policy quality and generalization. The problem is formulated as a partially observable Markov decision process (POMDP). A teacher model combining a dynamic graph attention network (GAT) and a large language model (LLM) is designed, where the GAT constructs a dynamic heterogeneous topology of UAVs, obstacles, and targets, while the LLM performs semantic reasoning to generate coordinated expert strategies. A joint action evaluation mechanism further builds a high-value expert knowledge base. The student model adopts a hybrid imitation–reinforcement framework. Imitation learning (IL) first clones expert behaviors from the knowledge base to accelerate policy acquisition, after which a centralized training with decentralized execution (CTDE)-based PPO stage refines the policy using interaction data, improving adaptability to wind disturbances and dynamic obstacles. This unified framework enables an effective transition from expert-guided learning to autonomous optimization. Simulation results show that LLM-PD substantially outperforms state-of-the-art baselines in path efficiency, convergence speed, cooperative stability, and generalization. Ablation studies further confirm the complementary benefits of LLM reasoning, GAT structural perception, and the imitation–reinforcement integration mechanism.
Keywords:
Imitation learning (IL)
Industrial Internet of Things (IIoT)
knowledge distillation
large language model (LLM)
multi-uncrewed aerial vehicle (UAV) systems
reinforcement learning (RL)

Journal

IEEE Internet of Things Journal cover
IEEE Internet of Things Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

Organization

C
chongqing university of posts and telecommunications
Scholars:
429
Papers: 176
Citations: 0