Return
RoLiC: A Robust LiDAR-Camera Fusion Framework for 3D Object Detection
DOI:10.1109/tip.2026.3705187.png)
Abstract
En 中文
In the 3D object detection task of autonomous driving systems, LiDAR and camera are the most crucial sensors, and current methods primarily focus on fusion strategies for these two complementary modalities. However, in real-world driving scenarios, potential sensor failures pose critical risks, potentially undermining the effectiveness of fusion-based detection approaches. This work presents RoLiC, a robust LiDAR–camera fusion framework designed to handle three challenging deployment scenarios within a unified model: LiDAR failure, camera failure, and simultaneous LiDAR–camera failure. To mitigate cross-modal dependency and recover missing information, RoLiC introduces two cross-modality feature transformers (L2C and C2L) that bidirectionally complete features between modalities when partial data is available. To further enhance feature reliability, we design a Sparse Similarity Loss (SSL) that constrains feature learning within high-probability object regions. Moreover, RoLiC integrates a task-aware two-stage feature knowledge distillation strategy (MS1 and MS2), where MS1 captures cross-modality complementarities, and MS2 distills knowledge between the fused modality and complete modality settings. Extensive experiments on the nuScenes and KITTI benchmarks demonstrate that RoLiC consistently outperforms state-of-the-art methods across all sensor-failure conditions. Code is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/JLIN77/RoLiC</uri>
Keywords:
Autonomous driving
sensor failures
robustness
feature transformer
feature completion
Journal
IF:
13.7
Papers:
1.0W
Citations:
8.4W

