Return
DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection
M
J
张
DOI:10.1109/tmm.2026.3668596.png)
Abstract
En 中文
State-of-the-art LiDAR-camera 3D object detectors usually focus on feature fusion, but often neglect the role of depth in modulating the contribution of different modalities. In this work, we observe that the importance of each modality varies with depth through statistical analysis and visualization. Motivated by this finding, we propose a Depth-Aware Hybrid Feature Fusion (DepthFusion) strategy that explicitly encodes metric depth into the BEV feature space, guiding the fusion weights of point cloud and RGB image features at both global and local levels. Specifically, the Depth-GFusion module adaptively adjusts the weights of image BEV features in multi-modal global features via explicit depth encoding. Furthermore, to recover information lost when projecting raw features to BEV space, the Depth-LFusion module dynamically adjusts the weights of voxel and multi-view image features in local features based on depth. Extensive experiments on the nuScenes and KITTI datasets demonstrate that DepthFusion outperforms previous state-of-the-art methods. Moreover, DepthFusion shows superior robustness to various corruptions on the nuScenes-C dataset.
Keywords:
Depth Encoding
Hybrid Feature Fusion
LiDAR-Camera 3D Object Detection
Journal
IF:
9.7
Papers:
4.4K
Citations:
2.4W
