arrow
Return

VisualRAG: Knowledge-Guided Retrieval Augmentation for Image-Text Matching

delete2025-08-08
delete0
PRE
AI
H
Hengchang Wang
刘丽 (Li Liu)
H
Huaxiang Zhang
L
Lei Zhu
X
Xiaojun Chang
杜浩 (Hao Du)
DOI:10.1109/TCSVT.2025.3597097delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image-text matching as a fundamental cross-modal understanding task presents unique challenges in weakly-aligned scenarios. Such data typically feature highly abstract textual captions with sparse entity References, creating a significant semantic gap with visual content. Current mainstream methods, primarily designed for strongly aligned data pairs, employ dynamic modeling or multi-dimensional similarity computation to achieve feature space mapping. However, they struggle with information asymmetry and modal heterogeneity in weakly aligned cases. To address this, we propose a Visual Perception Knowledge Enhancement (VPKE) framework. Unlike existing methods based on strong alignment assumptions, this framework mines latent image semantics through vision-language models and generates auxiliary captions, overcoming the information bottleneck of traditional text modalities. Its core innovation lies in an adaptive knowledge distillation mechanism that combines retrieval-augmented generation (RAG) with key entity extraction. This mechanism effectively filters noise when introducing external knowledge while optimizing cross-modal feature integration. The framework employs multi-level similarity evaluation to dynamically adjust fusion weights among original text, key entities, and auxiliary captions, enabling adaptive integration of diverse semantic features and significantly improving model flexibility. Additionally, multi-scale feature extraction further enhances cross-modal representation capabilities. Experimental results show that the proposed method performs excellently in image-text retrieval tasks on the MSCOCO and Flickr30K datasets, validating its effectiveness.
Keywords:
Image-text matching
knowledge enhancement
large language model
modality heterogeneity

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

T
tongji university
Scholars:
7.6W
Papers: 5.9W
Citations: 98
S
shandong normal university
Scholars:
1.0W
Papers: 8.2K
Citations: 3
U
University of Science and Technology of China
Scholars:
1.5W
Papers: 5.4K
Citations: 11.3W
researcher View more organizations