arrow
Return

Rwkv-vg: visual grounding with RWKV-driven encoder-decoder framework

delete2025-02-21
delete0
PRE
AI
F
Fudong Nian
Y
Yanhong Gu
W
Wentao Wang
A
Aoyu Liu
张东 (Dong Zhang)
L
Li, Fanding *
DOI:10.1007/s00530-025-01720-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual grounding is a fundamental task that bridges vision and language, aiming to accurately associate natural language queries with specific regions in an image. Existing approaches, predominantly based on Transformers or CNNs, struggle with balancing computational efficiency and fine-grained semantic alignment. In this paper, we propose RWKV-VG, the first visual grounding framework entirely built on the RWKV architecture. Leveraging RWKV's unique ability to combine RNN-like sequential modeling with Transformer-like attention, our model efficiently achieves both intra-modal and cross-modal reasoning. The framework consists of a RWKV-based visual encoder, a RWKV-based linguistic encoder, and a RWKV-based visual-linguistic decoder, complemented by a learnable [REG] token designed for box regression. Comprehensive evaluations on benchmark datasets, including ReferItGame and the RefCOCO series, demonstrate the superiority of RWKV-VG, achieving state-of-the-art performance with rapid convergence. Ablation studies further confirm the effectiveness of the RWKV modules and the [REG] token design. Our work establishes RWKV as a compelling alternative to conventional architectures for visual grounding tasks. To facilitate future research, the code and pre-trained models are released at https://github.com/nianfd/RWKV-VG.
Keywords:
Visual grounding
RWKV
Encoder-decoder
Cross-modal learning

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

J
Jiangnan University
Scholars:
3.9W
Papers: 2.7W
Citations: 4.7W
H
hefei university
Scholars:
2.2K
Papers: 1.3K
Citations: 20
J
Jilin University
Scholars:
8.6W
Papers: 5.5W
Citations: 8.9K
researcher View more organizations