Return
Fusing depth embeddings for monocular 3D object detection
DOI:10.1016/j.jvcir.2025.104555.png)
Abstract
En 中文
• A novel depth-assisted monocular 3D object detection method is proposed. The depth map, generated by a pre-trained monocular depth estimator, is embedded into a transformer-based detection framework as learnable depth embeddings, providing effective guidance for 3D object detection. • A lightweight keypoint detection module is designed to predict the 3D projective centers of objects. The features at these locations are then sampled to initialize object queries, effectively reducing the learning complexity of the transformer decoder.
Keywords:
monocular 3D object detection
depth estimation
transformer-based framework
keypoint detection
object query initialization
Journal
IF:
3.1
Papers:
414
Citations:
5.6K
Organization
No organization information available

