arrow
返回

Deep Relation Embedding for Cross-Modal Retrieval

delete2021-01-01
delete32
PRE
AI
Y
Yifan Zhang
W
Wengang Zhou
M
Min Wang
Q
Qi Tian
李
李厚强 (Houqiang Li) *
DOI:10.1109/TIP.2020.3038354delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Cross-modal retrieval aims to identify relevant data across different modalities. In this work, we are dedicated to cross-modal retrieval between images and text sentences, which is formulated into similarity measurement for each image-text pair. To this end, we propose a Cross-modal Relation Guided Network (CRGN) to embed image and text into a latent feature space. The CRGN model uses GRU to extract text feature and ResNet model to learn the globally guided image feature. Based on the global feature guiding and sentence generation learning, the relation between image regions can be modeled. The final image embedding is generated by a relation embedding module with an attention mechanism. With the image embeddings and text embeddings, we conduct cross-modal retrieval based on the cosine similarity. The learned embedding space well captures the inherent relevance between image and text. We evaluate our approach with extensive experiments on two public benchmark datasets, i.e., MS-COCO and Flickr30K. Experimental results demonstrate that our approach achieves better or comparable performance with the state-of-the-art methods with notable efficiency.
Keyword:
Semantics
Feature extraction
Visualization
Computational modeling
Task analysis
Training
Optimization
Image-text matching
cross-modal
retrieval
relation
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

H
huawei technologies
学者数:
3.3K
论文数: 2.9K
被引数: 1
U
university of science & technology of china, cas
学者数:
3.2W
论文数: 2.7W
被引数: 74
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

Learning Discriminative Binary Codes for Large-scale Cross-modal Retrieval
err2017-05-01
err382
PREAI
errXu, Xing; Shen, Fumin; Yang, Yang; Shen, Heng Tao; Li, Xuelong
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
err分享
err收藏
Ridged Fields in British Honduras
err1975-01-01
err0
PREAI
errG. W. Olson; A. H. Siemens; D. E. Puleston; G. Cal; D. Jenkins
err分享
err收藏
Probing vaccine antigens against bovine mastitis caused by Streptococcus uberis
err2016-07-01
err0
PREAI
errRosa Collado; Antoni Prenafeta; Luis González-González; Josep Antoni Pérez-Pons; Marta Sitjà
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
Semi-Paired Discrete Hashing: Learning Latent Hash Codes for Semi-Paired Cross-View Retrieval
err2017-12-01
err114
PREAI
errShen, Xiaobo; Shen, Fumin; Sun, Quan-Sen; Yang, Yang; Yuan, Yun-Hao; Shen, Heng Tao
err分享
err收藏
学者 查看更多内容