arrow
Return

Learning Relationship-Enhanced Semantic Graph for Fine-Grained Image-Text Matching

delete2024-02-01
delete7
PRE
AI
X
Xin Liu
Y
Yi He
Y
Yiu‐ming Cheung *
X
Xing Xu
王南南 cover
王南南 (Nannan Wang)
DOI:10.1109/TCYB.2022.3179020delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image-text matching of natural scenes has been a popular research topic in both computer vision and natural language processing communities. Recently, fine-grained image-text matching has shown its significant advance in inferring the high-level semantic correspondence by aggregating pairwise region-word similarity, but it remains challenging mainly due to insufficient representation of high-order semantic concepts and their explicit connections in one modality as its matched in another modality. To tackle this issue, we propose a relationship-enhanced semantic graph (ReSG) model, which can improve the image-text representations by learning their locally discriminative semantic concepts and then organizing their relationships in a contextual order. To be specific, two tailored graph encoders, visual relationship-enhanced graph (VReG) and textual relationship-enhanced graph (TReG), are respectively exploited to encode the high-level semantic concepts of corresponding instances and their semantic relationships. Meanwhile, the representations of each graph node are optimized by aggregating semantically contextual information to enhance the node-level semantic correspondence. Further, the hard-negative triplet ranking loss, center hinge loss, and positive-negative margin loss are jointly leveraged to learn the fine-grained correspondence between the ReSG representations of image and text, whereby the discriminative cross-modal embeddings can be explicitly obtained to benefit various image-text matching tasks in a more interpretable way. Extensive experiments verify the advantages of the proposed fine-grained graph matching approach, by achieving the state-of-the-art image-text matching results on public benchmark datasets.
Keywords:
Contextual information
high-level semantic concept
image-text matching
relationship-enhanced graph

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

H
Hong Kong Baptist University
Scholars:
6.3K
Papers: 7.5K
Citations: 1.3W
H
huaqiao university
Scholars:
1.1W
Papers: 7.1K
Citations: 131
X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K
researcher View more organizations