1
Return

DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection

delete2026-03-09
delete0
PRE
AI
M
Mingqian Ji
J
Jian Yang
张姗姗 (Shanshan Zhang)
DOI:10.1109/tmm.2026.3668596delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
State-of-the-art LiDAR-camera 3D object detectors usually focus on feature fusion, but often neglect the role of depth in modulating the contribution of different modalities. In this work, we observe that the importance of each modality varies with depth through statistical analysis and visualization. Motivated by this finding, we propose a Depth-Aware Hybrid Feature Fusion (DepthFusion) strategy that explicitly encodes metric depth into the BEV feature space, guiding the fusion weights of point cloud and RGB image features at both global and local levels. Specifically, the Depth-GFusion module adaptively adjusts the weights of image BEV features in multi-modal global features via explicit depth encoding. Furthermore, to recover information lost when projecting raw features to BEV space, the Depth-LFusion module dynamically adjusts the weights of voxel and multi-view image features in local features based on depth. Extensive experiments on the nuScenes and KITTI datasets demonstrate that DepthFusion outperforms previous state-of-the-art methods. Moreover, DepthFusion shows superior robustness to various corruptions on the nuScenes-C dataset.
Keywords:
Depth Encoding
Hybrid Feature Fusion
LiDAR-Camera 3D Object Detection

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.4K
Citations:
2.4W

Organization

N
nanjing university of science and technology
Scholars:
2.9K
Papers: 983
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers