Return
LM2CNet: Enhancing Monocular 3D Visual Grounding with Language Guided Multi-Modality Coupling Network
M
Q
S
J
L
G
DOI:10.1016/j.patrec.2026.03.023.png)
Abstract
En 中文
• We introduce a novel Mono3DRefer-nuScenes dataset tailored for the monocular 3D visual grounding task. • Our dataset accounts for diversity in both scenarios and languages, encompassing three language types. • We propose a novel multi-modality coupling network designed to utilize linguistic inputs to bridge the gap between visual and depth modalities • We evaluated various methods on the Mono3DRefer-nuScenes dataset, accompanied by visualization experiments.
Keywords:
Mono3DRefer-nuScenes
monocular 3D visual grounding
multi-modality coupling network
language-guided
depth modality
Journal
IF:
3.3
Papers:
7.8K
Citations:
1.6W
