arrow
返回

Learning Click-Based Deep Structure-Preserving Embeddings with Visual Attention

delete2019-08-08
delete3
PRE
AI
Y
Yehao Li
Y
Yingwei Pan
T
Ting Yao *
H
Hongyang Chao
Y
Yong Rui
T
Tao Mei
DOI:10.1145/3328994delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
One fundamental problem in image search is to learn the ranking functions (i.e., the similarity between query and image). Recent progress on this topic has evolved through two paradigms: the text-based model and image ranker learning. The former relies on image surrounding texts, making the similarity sensitive to the quality of textual descriptions. The latter may suffer from the robustness problem when human-labeled query-image pairs cannot represent user search intent precisely. We demonstrate in this article that the preceding two limitations can be well mitigated by learning a cross-view embedding that leverages click data. Specifically, a novel click-based Deep Structure-Preserving Embeddings with visual Attention (DSPEA) model is presented, which consists of two components: deep convolutional neural networks followed by image embedding layers for learning visual embedding, and a deep neural networks for generating query semantic embedding. Meanwhile, visual attention is incorporated at the top of the convolutional neural network to reflect the relevant regions of the image to the query. Furthermore, considering the high dimension of the query space, a new click-based representation on a query set is proposed for alleviating this sparsity problem. The whole network is end-to-end trained by optimizing a large margin objective that combines cross-view ranking constraints with in-view neighborhood structure preservation constraints. On a large-scale click-based image dataset with 11.7 million queries and 1 million images, our model is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods and achieves, to date, the best reported NDCG@25 of 52.21%.
Keyword:
Cross-view embedding
image search
click data
CNN
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

L
lenovo
学者数:
131
论文数: 110
被引数: 1
S
Sun Yat Sen University
学者数:
9.9W
论文数: 7.2W
被引数: 95
L
legend holdings
学者数:
208
论文数: 177
被引数: 1
学者 查看更多机构