Return
Advancing Multi-Robot Networks via MLLM-Driven Sensing, Communication, and Computation: A Comprehensive Survey
H
H
K
J
H
D
K
S
J
Y
C
M
B
DOI:10.1109/comst.2026.3683120.png)
Abstract
En 中文
Imagine a near future where advanced humanoid robots, powered by multimodal large language models (MLLMs), effortlessly interpret real-time sensing data, not just from their own sensors, but also from neighboring drones, autonomous vehicles, or underwater vehicles. Such robots for the physical AI are already taking shape in laboratories, hinting at imminent deployment across industries to tackle tasks like warehouse logistics, manufacturing, factory assembly, precision agriculture, public-space assistance, and on-site medical or safety rescue. While a single robot can demonstrate impressive local autonomy, realistic missions demand holistic coordination among multiple agents, compelling them to share and jointly interpret vast streams of sensor information. In these high-stakes scenarios, communication is indispensable, since without robust links, each robot remains blind to the broader mission context and cannot leverage the combined intelligence offered by a collective MLLM. Yet transmitting comprehensive sensor data from dozens or hundreds of robots, each with its own bandwidth and latency constraints, can overwhelm networks. This challenge is exacerbated when a system-level orchestrator or cloud-based MLLM needs to fuse multimodal inputs to generate holistic decisions, like route planning, anomaly detection, or real-time adjustments to complex collaborative tasks. Crucially, these tasks are often initiated by high-level natural language instructions (e.g., “Search for the yellow bin”). This text-based intent serves as a powerful filter for resource optimization: by understanding the specific goal via MLLMs, the system can selectively activate only the relevant sensing modalities, dynamically allocate communication bandwidth, and determine the optimal computation placement. This “intent-to-resource” mapping capability is a fundamental motivation for the proposed unified architecture. Moreover, many real deployments require open-vocabulary perception and language-grounded action (e.g., recognizing previously unseen objects or infrastructure outside the robot’s field-of-view), which is difficult to achieve with closed-set on-device perception alone. Viewed this way, R2X is fundamentally an intent-to-resource orchestration problem: given a high-level language command and system context, the network must jointly optimize sensing, wireless communication, and computation so that task-level success is maximized under resource constraints. This survey examines how integrated sensing, communication, and computation design paves the way for effective multi-robot coordination under MLLM guidance. We begin by reviewing state-of-the-art sensing modalities (from bounding-box LiDAR to hyperspectral imaging) and their interplay with semantic or partial-compression techniques. Next, we analyze communication strategies, including low latency protocols, edge-based orchestration, and adaptive resource allocation, to ensure scalable and timely data delivery at scale. We then explore distributed, hybrid, and fully centralized computing approaches, highlighting how large-model reasoning can be split among on-device distillation and powerful edge/cloud servers. To ground these concepts in measurable outcomes, we further present four end-to-end demonstrations that make the orchestration loop explicit (sense <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\rightarrow $ </tex-math></inline-formula> communicate <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\rightarrow $ </tex-math></inline-formula> compute <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\rightarrow $ </tex-math></inline-formula> act): 1) digital-twin warehouse navigation with semantic sensing and predictive link context; 2) mobility-driven proactive MCS control under delayed feedback; 3) a real FollowMe robot with a practical semantic-sensing switch over WiFi; and 4) real-hardware open-vocabulary trash sorting where an edge-assisted MLLM grounds text instructions to unseen objects and out-of-FOV bins via multi-view sensing. Ultimately, our goal is to guide the robotics research community and industry practitioners in devising end-to-end solutions that integrate sensing, communication, and computation for real-time and large-scale multi-robot operations. By uniting advanced sensors, flexible network design, and MLLM-based intelligence, autonomous robot teams can achieve a truly holistic awareness and responsiveness in the complex environments of tomorrow. Across the demonstrations, we emphasize system-level metrics—payload size, end-to-end latency, reliability, and task success—to clarify when and why R2X-style edge/communication-assisted orchestration provides tangible advantages over purely on-device baselines.
Keywords:
Multimodal large language model (MLLM)
robot-to-everything (R2X)
autonomous communications
sensing/communication/computation-integrated system
Journal
I
IF:
46.7
Papers:
1.5K
Citations:
3.3W
