Return
Reliable Task Allocation for Multi-Robot System Using Safe Deep Reinforcement Learning
DOI:10.1109/tr.2026.3727234.png)
Abstract
En 中文
This paper addresses the critical reliability and safety challenges in multi-station multi-robot welding task allocation for automated manufacturing, where conventional task assignment approaches often overlook physical collision risks, leading to system failures and production downtime. Unlike existing methods that either ignore safety constraints or incorporate them heuristically as penalty terms in reward functions, we formulate the multi-robot task allocation problem as a Constrained Markov Decision Process (CMDP), enabling explicit and rigorous enforcement of operational safety constraints. Within this framework, we develop a safe deep reinforcement learning architecture that integrates a graph-based encoder for spatial feature extraction with a Lagrangian-based constrained policy optimization mechanism. This approach systematically balances the dual objectives of minimizing production cycle time and ensuring collision-free robot operations through adaptive safety-weight adjustment during training. Extensive experiments on both real-world welding cases and large-scale synthetic scenarios demonstrate that our method achieves zero safety violations while maintaining superior task efficiency, outperforming classical heuristics and conventional reinforcement learning baselines that exhibit notable collision rates. The results underscore the effectiveness of embedding safety as a formal constraint rather than an afterthought, providing a reliable and practical solution for industrial multi-robot coordination.
Keywords:
Industrial robot
task allocation
deep reinforcement learning
safe reinforcement learning
Constrained Markov Decision Process
Journal
IF:
5.7
Papers:
2.7K
Citations:
8.5K

