arrow
Return

MLA3DVG: Dual-modality multi-level semantic alignment for robust monocular 3D visual grounding

delete2026-08-17
delete0
PRE
AI
S
Shidi Chen
Z
Zheming Xu
L
Lili Wei
Z
Ziyi Chen
G
Gang Wen
郎丛妍 (Congyan Lang) *
B
Binyang Song
DOI:10.1016/j.patcog.2026.114622delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• MLA3DVG for language-guided 3D object localization from a single RGB image. • Multi-level decomposition models scene-, attribute-, and spatial-level semantics. • Multi-Level semantic alignment for fine-grained, level-wise cross-modal alignment. • A spatial-level semantic reconstruction loss enhances geometry-aware understanding. • Notable improvements in long-range and ambiguous object grounding. Abstract Monocular 3D visual grounding (Mono3DVG) aims to localize objects in 3D space from natural language descriptions using only RGB images, offering a more scalable alternative to prior approaches that require expensive depth sensors. However, the absence of explicit geometric cues hinders spatial reasoning, and existing approaches typically perform coarse scene-level semantic-geometric alignment, overlooking the hierarchical linguistic structure that encodes fine-grained attributes and spatial relations for precise localization. To address this limitation, we propose MLA3DVG, a dual-modality framework for 3D visual grounding that explicitly models and aligns multi-level semantics across visual and textual modalities. Specifically, a Multi-Level Semantic Decomposition (MLSD) module decomposes both textual and visual inputs into scene-, attribute-, and spatial-level representations, capturing fine-grained geometric features for richer spatial understanding. These representations are subsequently integrated through a Multi-Level Semantic Alignment (MLSA) mechanism, which enables fine-grained, level-wise cross-modal alignment. Furthermore, we introduce a Semantic-Guided Grounding Decoder (SGGD) with dual-modality supervision, which enhances geometry-aware reasoning by reconstructing masked spatial expressions within the textual domain. Extensive experiments on the Mono3DRefer benchmark demonstrate that MLA3DVG achieves state-of-the-art performance, with notable improvements in challenging scenarios involving long-range targets and multiple similar objects. The source code will be available at MLA3DVG.
Keywords:
Monocular 3D visual grounding
Vision-language understanding
Cross-modal alignment

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

B
Beijing Jiaotong University
Scholars:
2.2W
Papers: 1.7W
Citations: 1.2W
N
Northwest Normal University
Scholars:
988
Papers: 249
Citations: 5.8K
N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
researcher View more organizations