arrow
返回

Visual context learning based on textual knowledge for image-text retrieval

delete2022-08-01
delete7
PRE
AI
Y
Yuzhuo Qin
X
Xiaodong Gu *
Z
Zhenshan Tan
DOI:10.1016/j.neunet.2022.05.008delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Image-text bidirectional retrieval is a significant task within cross-modal learning field. The main issue lies on the jointly embedding learning and accurately measuring image-text matching score. Most prior works make use of either intra-modality methods performing within two separate modalities or inter-modality ones combining two modalities tightly. However, intra-modality methods remain ambiguous when learning visual context due to the existence of redundant messages. And inter-modality methods increase the complexity of retrieval because of unifying two modalities closely when learning modal features. In this research, we propose an eclectic Visual Context Learning based on Textual knowledge Network (VCLTN), which transfers textual knowledge to visual modality for context learning and decreases the discrepancy of information capacity between two modalities. Specifically, VCLTN merges label semantics into corresponding regional features and employs those labels as intermediaries between images and texts for better modal alignment. Contextual knowledge of those labels learned within textual modality is utilized to guide the visual context learning. Besides, considering the homogeneity within each modality, global features are merged into regional features for assisting in the context learning. In order to alleviate the imbalance of information capacity between images and texts, entities together with relations inside the given caption are extracted and an auxiliary caption is sampled for attaching supplementary messages to textual modality. Experiments performed on Flickr30K and MS-COCO reveal that our model VCLTN achieves best results compared with the state-of-the-art methods. (C) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Image-text retrieval
Knowledge transfer
Visual context learning
Modal alignment

期刊

Neural Networks 封面图
Neural Networks
IF:
6.3
论文数:
8.2K
被引数:
3.0W

机构

F
fudan university
学者数:
11.8W
论文数: 7.7W
被引数: 121
引用论文

引用论文

err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Parameter-Free Model of the Self-Catalyzed Growth of Ga(As,P) Nanowires
err2022-05-11
err0
PREAI
errN. V. Sibirev; Yu. S. Berdnikov; V. V. Fedorov; I. V. Shtrom; A. D. Bolshakov
err分享
err收藏
Bidirectional image-sentence retrieval by local and global deep matching
err2019-06-01
err29
PREAI
errMa, Lin; Jiang, Wenhao; Jie, Zequn; Wang, Xu
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Using collective system design to define and communicate organization goals and related solutions
err2017-10-12
err0
PREAI
errDavid Cochran; Gordon Schmidt; Jennifer Oxtoby; Mike Hensley; Jason Barnes
err分享
err收藏
学者 查看更多内容