arrow
Return

CLMPO-EC: A Lightweight Multi-UAV Multiarea Coverage Path Planning Method Using Deep Reinforcement Learning

delete2026-02-03
delete0
PRE
AI
Z
Zhichao Qian
冯勇 cover
冯勇 (Yong Feng)
N
Nianbo Liu
Q
Qian Qian
DOI:10.1109/JIOT.2026.3659864delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Unmanned aerial vehicles (UAVs), as an aerial extension of the Internet of Things (IoT) sensing layer, have played an increasingly important role in applications such as environmental monitoring, disaster assessment, and precision agriculture. These tasks can be uniformly abstracted as coverage path planning (CPP), which aims to achieve efficient scanning and surveying while ensuring complete coverage of single or multiple disconnected regions. While single-region CPP has been extensively studied, in multiregion settings existing methods often rely on predefined coverage patterns to guarantee completeness, which to some extent limits their flexibility. Meanwhile, constraints on UAV energy and onboard computation impose higher performance requirements on planning methods. To address these challenges, this article targets an energy-constrained multi-UAV cooperative scenario and proposes a cross-layer, energy-constrained path optimization framework based on multiagent reinforcement learning (CLMPO-EC). Specifically, the framework organizes the overall task into two layers—CPP and multiagent path planning (MAPP)—and, on this basis, integrates back-and-forth path planning (BFP) with multiagent reinforcement learning (MARL) under a centralized training and distributed execution paradigm to construct a unified, interactive, and structured environmental model. CLMPO-EC further introduces a lightweight cross-layer connection network that propagates raw state information to higher layers to enhance learning efficiency. In addition, building on BFP, an entrance—exit exploration factor is proposed to dynamically adjust the exploration probability of regional entrances and exits in CPP according to the training phase and batch, thereby improving the efficiency of searching for optimal solutions. Theoretical analysis and experimental results demonstrate that the proposed method achieves superior performance in terms of optimality and efficiency.
Keywords:
Coverage path planning (CPP)
deep reinforcement learning (DRL)
Internet of Things (IoT)
multiple regions
multiple unmanned aerial vehicles (UAVs)

Journal

IEEE Internet of Things Journal cover
IEEE Internet of Things Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

Organization

U
university of electronic science and technology of china
Scholars:
1.2W
Papers: 4.6K
Citations: 4
K
kunming university of science and technology
Scholars:
5.0K
Papers: 1.4K
Citations: 0