arrow
返回

Learning Relationship-Enhanced Semantic Graph for Fine-Grained Image-Text Matching

delete2024-02-01
delete7
PRE
AI
X
Xin Liu
Y
Yi He
Y
Yiu‐ming Cheung *
X
Xing Xu
王南南 封面图
王南南 (Nannan Wang)
DOI:10.1109/TCYB.2022.3179020delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Image-text matching of natural scenes has been a popular research topic in both computer vision and natural language processing communities. Recently, fine-grained image-text matching has shown its significant advance in inferring the high-level semantic correspondence by aggregating pairwise region-word similarity, but it remains challenging mainly due to insufficient representation of high-order semantic concepts and their explicit connections in one modality as its matched in another modality. To tackle this issue, we propose a relationship-enhanced semantic graph (ReSG) model, which can improve the image-text representations by learning their locally discriminative semantic concepts and then organizing their relationships in a contextual order. To be specific, two tailored graph encoders, visual relationship-enhanced graph (VReG) and textual relationship-enhanced graph (TReG), are respectively exploited to encode the high-level semantic concepts of corresponding instances and their semantic relationships. Meanwhile, the representations of each graph node are optimized by aggregating semantically contextual information to enhance the node-level semantic correspondence. Further, the hard-negative triplet ranking loss, center hinge loss, and positive-negative margin loss are jointly leveraged to learn the fine-grained correspondence between the ReSG representations of image and text, whereby the discriminative cross-modal embeddings can be explicitly obtained to benefit various image-text matching tasks in a more interpretable way. Extensive experiments verify the advantages of the proposed fine-grained graph matching approach, by achieving the state-of-the-art image-text matching results on public benchmark datasets.
Keyword:
Contextual information
high-level semantic concept
image-text matching
relationship-enhanced graph

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

H
Hong Kong Baptist University
学者数:
6.3K
论文数: 7.5K
被引数: 1.3W
H
huaqiao university
学者数:
1.1W
论文数: 7.1K
被引数: 131
X
Xidian University
学者数:
2.4W
论文数: 1.9W
被引数: 9.7K
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Parameter-Free Model of the Self-Catalyzed Growth of Ga(As,P) Nanowires
err2022-05-11
err0
PREAI
errN. V. Sibirev; Yu. S. Berdnikov; V. V. Fedorov; I. V. Shtrom; A. D. Bolshakov
err分享
err收藏
Evaluation of Antioxidant and Free Radical Scavenging Capacities of Polyphenolics from Pods of Caesalpinia pulcherrima
err2012-05-18
err0
errOAAI
errFeng-Lin Hsu; Wei-Jan Huang; Tzu-Hua Wu; Mei-Hsien Lee; Lih-Chi Chen; Hsiao-Jen Lu; Wen-Chi Hou; Mei-Hsiang Lin
err分享
err收藏
学者 查看更多内容