Return
Decoding before aligning: Scale-Adaptive Early-Decoding Transformer for visual grounding
DOI:10.1016/j.neucom.2025.129756.png)
Abstract
En 中文
Visual grounding, the task of precisely localizing objects within an image as described by natural language expressions, has recently seen significant advancements through the integration of transformer decoders and early-fusion mechanisms. These approaches aim to enhance the decoding of crucial grounding-related semantics and the generation of discriminative representations. However, they display significant limitations, including reduced generalizability due to the need for custom plug-ins, a gap in effectively bridging decoded visual and linguistic features, and difficulties in processing single-scale features within complex, hierarchical scenes. To address these issues, this study introduces the Scale-Adaptive Early-Decoding Transformer (SA-EDTR), a novel approach that combines decoding and early-fusion mechanisms within a single transformer layer. By positioning the decoding stage before the alignment stage, the decoder adaptively generates richer, context- aware representations that guide the aligning module. This design allows SA-EDTR to effectively overcome previous limitations in feature alignment. Additionally, we integrate a Scale-Adaptive Fusion network (SAF) to fuse multi-scale and heterogeneous representations from parallel Early-Decoding Transformer layers, enhancing the capture of spatial, coarse, and fine-grained semantic details. Extensive experiments demonstrate that SAEDTR outperforms current state-of-the-art methods in both REC and RES tasks across various mainstream visual grounding datasets, including RefCOCO, RefCOCO+, RefCOCOg-umd, and RefCOCOg-google. The results highlight the architecture's remarkable capability inaccurately localizing objects and interpreting intricate scenes.
Keywords:
Referring expression comprehension
Referring expression segmentation
Visual grounding
Vision-language
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

