Return
Transformer-based multi-modal feature fusion for end-to-end autonomous driving
X
H
Y
DOI:10.1016/j.ijtst.2025.10.017.png)
Abstract
En 中文
End-to-end autonomous driving has drawn significant attention for its ability to unify perception, decision-making, and control into a single learning framework. However, existing methods often struggle in complex and dynamic environments due to ineffective multi-modal feature fusion and inflexible decision-making strategies. To overcome these challenges, we proposed a novel Transformer-based multi-modal feature fusion framework that integrates red-green-blue (RGB) images and depth data to generate robust vehicular control commands. Our approach employed a four-stage pyramid vision transformer (PVT) backbone to extract multi-scale features and introduced a dual-attention feature fusion network to capture both intra-modal and cross-modal dependencies, yielding a more robust and context-aware environmental representation. Furthermore, we proposed a dynamic trajectory-control fusion strategy (Traj-CtrlFuser), which utilizes a learnable loss estimator to adaptively balance the outputs of trajectory and control branches based on real-time driving conditions. Extensive evaluations on the CARLA Town05 Long benchmark demonstrated that our model outperformed the current state-of-the-art multi-modal method, DriveInsight, achieving a 5.6% improvement in driving score and a 7.6% improvement in infraction score. These results underscore the framework’s potential to enhance the robustness and safety of autonomous vehicles in complex urban scenarios, providing a reliable and scalable solution for future end-to-end driving systems.
Keywords:
End-to-end autonomous driving
multi-modal fusion
Transformer
Adaptive control
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.8
Papers:
1.5K
Citations:
1.5K
