Return
MLA3DVG: Dual-modality multi-level semantic alignment for robust monocular 3D visual grounding
DOI:10.1016/j.patcog.2026.114622.png)
Abstract
En 中文
•
MLA3DVG for language-guided 3D object localization from a single RGB image.
•
Multi-level decomposition models scene-, attribute-, and spatial-level semantics.
•
Multi-Level semantic alignment for fine-grained, level-wise cross-modal alignment.
•
A spatial-level semantic reconstruction loss enhances geometry-aware understanding.
•
Notable improvements in long-range and ambiguous object grounding.
Abstract
Monocular 3D visual grounding (Mono3DVG) aims to localize objects in 3D space from natural language descriptions using only RGB images, offering a more scalable alternative to prior approaches that require expensive depth sensors. However, the absence of explicit geometric cues hinders spatial reasoning, and existing approaches typically perform coarse scene-level semantic-geometric alignment, overlooking the hierarchical linguistic structure that encodes fine-grained attributes and spatial relations for precise localization. To address this limitation, we propose MLA3DVG, a dual-modality framework for 3D visual grounding that explicitly models and aligns multi-level semantics across visual and textual modalities. Specifically, a Multi-Level Semantic Decomposition (MLSD) module decomposes both textual and visual inputs into scene-, attribute-, and spatial-level representations, capturing fine-grained geometric features for richer spatial understanding. These representations are subsequently integrated through a Multi-Level Semantic Alignment (MLSA) mechanism, which enables fine-grained, level-wise cross-modal alignment. Furthermore, we introduce a Semantic-Guided Grounding Decoder (SGGD) with dual-modality supervision, which enhances geometry-aware reasoning by reconstructing masked spatial expressions within the textual domain. Extensive experiments on the Mono3DRefer benchmark demonstrate that MLA3DVG achieves state-of-the-art performance, with notable improvements in challenging scenarios involving long-range targets and multiple similar objects. The source code will be available at MLA3DVG.
Keywords:
Monocular 3D visual grounding
Vision-language understanding
Cross-modal alignment
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

