返回
Multi-level semantics probability embedding for image-text matching
DOI:10.1016/j.ipm.2024.103968.png)
摘要
En 中文
The requirement of image-text matching is to retrieve matching images or texts based on textual or visual queries. However, image-text matching is inherently a many-to-many problem, as an image can correspond to multiple levels of visual semantic scenes, which can be described by different texts. Similarly, textual descriptions can be visualized through multiple visual scenes. This leads to ambiguity in the matching between images and texts. To better capture these matching relationships, we employ graph convolutional networks to extract multi-level semantic information for image-text pairs, and construct Gaussian distribution representations for image and text instead of conventional point representations. Furthermore, we introduce a inter-modal mixture of Gaussian distribution to constrain the matching relationships between image-text pairs, which ensures more precise distribution representations in a shared space and strengthens the correlation between cross-modal. We conducted experiments on Flickr30K and MS-COCO, which are two widely used datasets, demonstrates the superior performance of our approach.
Keyword:
Image-text matching
Probability embedding
Gaussian distribution
期刊
I
IF:
6.9
论文数:
5.2K
被引数:
1.4W
机构
引用论文
Global-local fusion based on adversarial sample generation for image-text matching
INFORMATION FUSION
IF15.5
603 VEDOLIZUMAB PREVENTS POSTOPERATIVE REUCURRENCE IN CROHN'S DISEASE: RESULTS OF THE REPREVIO TRIAL
MKVSE: Multimodal Knowledge Enhanced Visual-semantic Embedding for Image-text RetrievalMKVSE: 用于图像文本检索的多模态知识增强的视觉语义嵌入

