Return
MA-APD: Multiagent Asynchronous Probability-Decomposed Policy Gradient for Time-Constrained Moving Target Search
Q
H
C
D
DOI:10.1109/tro.2026.3710426.png)
Abstract
En 中文
This article investigates the time-constrained multirobot efficient search (MuRES) problem, focusing on the asynchronous coordination among multiple robots. To the best of authors’ knowledge, almost all MuRES solutions adopt the synchronous decision-making framework, wherein the robots act simultaneously. The synchronous assumption simplifies the MuRES problem formulation and facilitates the development of coordination search strategies. However, in real-world scenarios, spatial variability and heterogeneous motion characteristics render the synchronous execution scheme inefficient, leading to large idle time and poor task allocation. To address these limitations, we introduce the multiagent asynchronous probability-decomposed policy gradient (MA-APD) algorithm, which targets the asynchronous multirobot efficient search problem. MA-APD comprises two core components: first, the generalized value function approximation module, which evaluates and asynchronously decomposes the nonconvex time-constrained MuRES objective, and second, the asynchronous policy gradient module, which maps the decomposed value function onto individual robot’s decision-making timeline, enabling the asynchronous actor updates. To improve MA-APD’s sample efficiency and training stability, we further introduce an off-policy learning mechanism, a soft update module, and an entropy regularization term. Extensive simulation results in canonical MuRES environments indicate that MA-APD consistently outperforms canonical MuRES solutions as well as multiagent reinforcement learning algorithms. Furthermore, MA-APD is deployed to a physical multirobot system for moving target search in both self-constructed and realistic indoor environments, validating its practical effectiveness.
Keywords:
Multiagent reinforcement learning (MARL)
multirobot efficient search (MuRES)
time-constrained moving target search
Journal
IF:
10.5
Papers:
3.3K
Citations:
2.8W
