arrow
返回

Dual-path Convolutional Image-Text Embeddings with Instance Loss

delete2020-05-22
delete345
delete
OA
AI
Z
Zhedong Zheng *
L
Liang Zheng
M
Michael Garrett
Y
Yi Yang
X
Xu, Mingliang
S
Shen, Yi-Dong
DOI:10.1145/3383184delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Matching images and sentences demands a fine understanding of both modalities. In this article, we propose a new system to discriminatively embed the image and text to a shared visual-textual space. In this field, most existing works apply the ranking loss to pull the positive image/text pairs close and push the negative pairs apart from each other. However, directly deploying the ranking loss on heterogeneous features (i.e., text and image features) is less effective, because it is hard to find appropriate triplets at the beginning. So the naive way of using the ranking loss may compromise the network from learning inter-modal relationship. To address this problem, we propose the instance loss, which explicitly considers the intra-modal data distribution. It is based on an unsupervised assumption that each image/text group can be viewed as a class. So the network can learn the fine granularity from every image/text group. The experiment shows that the instance loss offers better weight initialization for the ranking loss, so that more discriminative embeddings can be learned. Besides, existing works usually apply the off-the-shelf features, i.e., word2vec and fixed visual feature. So in a minor contribution, this article constructs an end-to-end dual-path convolutional network to learn the image and text representations. End-to-end learning allows the system to directly learn from the data and fully utilize the supervision. On two generic retrieval datasets (Flickr30k and MSCOCO), experiments demonstrate that our method yields competitive accuracy compared to state-of-the-art methods. Moreover, in language-based person retrieval, we improve the state of the art by a large margin. The code has been made publicly available.
Keyword:
Image-sentence retrieval
cross-modal retrieval
language-based person search
convolutional neural networks
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

E
Edith Cowan University
学者数:
4.6K
论文数: 5.1K
被引数: 9.2K
I
institute of software, cas
学者数:
446
论文数: 388
被引数: 0
A
Australian National University
学者数:
2.1W
论文数: 2.3W
被引数: 3.9W
U
university of technology sydney
学者数:
1.6W
论文数: 2.0W
被引数: 25
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Flexible Multi-View Dimensionality Co-Reduction
err2017-02-01
err102
PREAI
errZhang, Changqing; Fu, Huazhu; Hu, Qinghua; Zhu, Pengfei; Cao, Xiaochun
err分享
err收藏
Twitter100k: A Real-World Dataset for Weakly Supervised Cross-Media Retrieval
err2018-04-01
err47
errOAAI
errHu, Yuting; Zheng, Liang; Yang, Yi; Huang, Yongfeng
err分享
err收藏
err分享
err收藏
学者 查看更多内容