返回
Combination subspace graph learning for cross-modal retrieval
DOI:10.1016/j.aej.2020.02.034.png)
摘要
En 中文
In this paper, a novel supervised cross-modal retrieval method, combination subspace graph learning (CSGL), is proposed, which primarily concentrates on the cross retrieval between images and texts, i.e., image retrieving texts (I2T) and text retrieving images (T2I). To project multi-modal data from a low-level feature space into a latent common subspace, most classical methods would learn the projective matrix separately for each mode, ignoring the consistency between different modes. Graph regularization is added to our objective function to preserve the structure of the original data in the projective space. Furthermore, to avoid the suboptimal solution during optimization, we use the collaborative learning strategy to obtain the projective matrix directly, which unites all the modes for better projection. Generally, the CSGL method takes the advantage of the semantic information and the original distribution of the image and text to obtain a more discriminative projection, which is learned in combination, rather than learning individually. Experimental results on three benchmark datasets, Wikipedia, Pascal Sentence, and INRIA-Web-search, show that the proposed method outperforms the state-of-the-art methods. (C) 2020 The Authors. Published by Elsevier B.V. on behalf of Faculty of Engineering,
Keyword:
Cross-modal retrieval
Collaborative learning
Graph regularization
Subspace learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.8
论文数:
6.3K
被引数:
2.6W
机构
引用论文
Learning unified binary codes for cross-modal retrieval via latent semantic hashing
NEUROCOMPUTING
IF6.5
CCL: Cross-modal Correlation Learning With Multigrained Fusion by Hierarchical NetworkCCL: 分层网络多粒度融合的跨模态相关学习
A numerical methodology for thermo-fluid dynamic modelling of tyre inner chamber: towards real time applications
Meccanica
IF0

