arrow
Return

Decoding before aligning: Scale-Adaptive Early-Decoding Transformer for visual grounding

delete2025-02-01
delete0
PRE
AI
L
Liuwu Li
蔡毅 cover
蔡毅 (Yi Cai)
王杰新 (Jiexin Wang) *
C
Cantao Wu
Q
Qingbao Huang
Q
Qing Li
DOI:10.1016/j.neucom.2025.129756delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual grounding, the task of precisely localizing objects within an image as described by natural language expressions, has recently seen significant advancements through the integration of transformer decoders and early-fusion mechanisms. These approaches aim to enhance the decoding of crucial grounding-related semantics and the generation of discriminative representations. However, they display significant limitations, including reduced generalizability due to the need for custom plug-ins, a gap in effectively bridging decoded visual and linguistic features, and difficulties in processing single-scale features within complex, hierarchical scenes. To address these issues, this study introduces the Scale-Adaptive Early-Decoding Transformer (SA-EDTR), a novel approach that combines decoding and early-fusion mechanisms within a single transformer layer. By positioning the decoding stage before the alignment stage, the decoder adaptively generates richer, context- aware representations that guide the aligning module. This design allows SA-EDTR to effectively overcome previous limitations in feature alignment. Additionally, we integrate a Scale-Adaptive Fusion network (SAF) to fuse multi-scale and heterogeneous representations from parallel Early-Decoding Transformer layers, enhancing the capture of spatial, coarse, and fine-grained semantic details. Extensive experiments demonstrate that SAEDTR outperforms current state-of-the-art methods in both REC and RES tasks across various mainstream visual grounding datasets, including RefCOCO, RefCOCO+, RefCOCOg-umd, and RefCOCOg-google. The results highlight the architecture's remarkable capability inaccurately localizing objects and interpreting intricate scenes.
Keywords:
Referring expression comprehension
Referring expression segmentation
Visual grounding
Vision-language

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

H
hong kong polytechnic university
Scholars:
3.0W
Papers: 4.1W
Citations: 921
S
south china university of technology
Scholars:
6.7W
Papers: 5.1W
Citations: 85
G
guangxi university
Scholars:
3.3W
Papers: 1.8W
Citations: 25
researcher View more organizations