arrow
Return

DeepInteraction++: Multi-Modality Interaction for Autonomous Driving

delete2025-08-01
delete0
delete
OA
AI
Z
Zeyu Yang
N
Nan Song
李炜 cover
李炜 (Wei Li)
X
Xiatian Zhu
L
Li Zhang
P
Philip H. S. Torr
DOI:10.1109/TPAMI.2025.3565194delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Existing top-performance autonomous driving systems typically rely on the <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">multi-modal fusion</i> strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and finally hampering the model performance. To address this limitation, in this work, we introduce a novel <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">modality interaction</i> strategy that allows individual per-modality representations to be learned and maintained throughout, enabling their unique characteristics to be exploited during the whole perception pipeline. To demonstrate the effectiveness of the proposed strategy, we design <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DeepInteraction++</i>, a multi-modal interaction framework characterized by a multi-modal representational interaction encoder and a multi-modal predictive interaction decoder. Specifically, the encoder is implemented as a dual-stream Transformer with specialized attention operation for information exchange and integration between separate modality-specific representations. Our multi-modal representational learning incorporates both object-centric, precise sampling-based feature alignment and global dense information spreading, essential for the more challenging planning task. The decoder is designed to iteratively refine the predictions by alternately aggregating information from separate representations in a unified modality-agnostic manner, realizing multi-modal predictive interaction. Extensive experiments demonstrate the superior performance of the proposed framework on both 3D object detection and end-to-end autonomous driving tasks.
Keywords:
Autonomous driving
3D object detection
multi-modal fusion

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

F
fudan university
Scholars:
11.6W
Papers: 7.7W
Citations: 121
N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
U
university of oxford
Scholars:
9.7W
Papers: 8.6W
Citations: 137
U
University of Surrey
Scholars:
1.2W
Papers: 1.3W
Citations: 22
researcher View more organizations