arrow
Return

Fusing depth embeddings for monocular 3D object detection

delete2025-08-07
delete0
PRE
AI
C
Chaofeng Ji *
赵丹 (Dan Zhao)
W
Wei Li
刘广恩 (Guangen Liu)
DOI:10.1016/j.jvcir.2025.104555delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• A novel depth-assisted monocular 3D object detection method is proposed. The depth map, generated by a pre-trained monocular depth estimator, is embedded into a transformer-based detection framework as learnable depth embeddings, providing effective guidance for 3D object detection. • A lightweight keypoint detection module is designed to predict the 3D projective centers of objects. The features at these locations are then sampled to initialize object queries, effectively reducing the learning complexity of the transformer decoder.
Keywords:
monocular 3D object detection
depth estimation
transformer-based framework
keypoint detection
object query initialization

Journal

Journal of Visual Communication and Image Representation cover
Journal of Visual Communication and Image Representation
IF:
3.1
Papers:
414
Citations:
5.6K

Organization

No organization information available