1
Return

MA-APD: Multiagent Asynchronous Probability-Decomposed Policy Gradient for Time-Constrained Moving Target Search

delete2026-07-06
delete0
PRE
AI
Q
Qihang Peng
H
Hongliang Guo
C
Chih‐Yung Wen
D
Daniela Rus
DOI:10.1109/tro.2026.3710426delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article investigates the time-constrained multirobot efficient search (MuRES) problem, focusing on the asynchronous coordination among multiple robots. To the best of authors’ knowledge, almost all MuRES solutions adopt the synchronous decision-making framework, wherein the robots act simultaneously. The synchronous assumption simplifies the MuRES problem formulation and facilitates the development of coordination search strategies. However, in real-world scenarios, spatial variability and heterogeneous motion characteristics render the synchronous execution scheme inefficient, leading to large idle time and poor task allocation. To address these limitations, we introduce the multiagent asynchronous probability-decomposed policy gradient (MA-APD) algorithm, which targets the asynchronous multirobot efficient search problem. MA-APD comprises two core components: first, the generalized value function approximation module, which evaluates and asynchronously decomposes the nonconvex time-constrained MuRES objective, and second, the asynchronous policy gradient module, which maps the decomposed value function onto individual robot’s decision-making timeline, enabling the asynchronous actor updates. To improve MA-APD’s sample efficiency and training stability, we further introduce an off-policy learning mechanism, a soft update module, and an entropy regularization term. Extensive simulation results in canonical MuRES environments indicate that MA-APD consistently outperforms canonical MuRES solutions as well as multiagent reinforcement learning algorithms. Furthermore, MA-APD is deployed to a physical multirobot system for moving target search in both self-constructed and realistic indoor environments, validating its practical effectiveness.
Keywords:
Multiagent reinforcement learning (MARL)
multirobot efficient search (MuRES)
time-constrained moving target search

Journal

IEEE Transactions on Robotics cover
IEEE Transactions on Robotics
IF:
10.5
Papers:
3.3K
Citations:
2.8W

Organization

T
the hong kong polytechnic university
Scholars:
3.9K
Papers: 2.3K
Citations: 0
A
agency for science, technology and research
Scholars:
502
Papers: 196
Citations: 0
M
massachusetts institute of technology
Scholars:
3.3K
Papers: 1.2K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers