arrow
返回

Learning Compact Hash Codes for Multimodal Representations Using Orthogonal Deep Structure

delete2015-09-01
delete84
PRE
AI
D
Daixin Wang *
崔鹏 封面图
崔鹏 (Peng Cui)
Wenwu Zhu 封面图
Wenwu Zhu (Wenwu Zhu)
DOI:10.1109/TMM.2015.2455415delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
As large-scale multimodal data are ubiquitous in many real-world applications, learning multimodal representations for efficient retrieval is a fundamental problem. Most existing methods adopt shallow structures to perform multimodal representation learning. Due to a limitation of learning ability of shallow structures, they fail to capture the correlation of multiple modalities. Recently, multimodal deep learning was proposed and had proven its superiority in representing multimodal data due to its high nonlinearity. However, in order to learn compact and accurate representations, how to reduce the redundant information lying in the multimodal representations and incorporate different complexities of different modalities in the deep models is still an open problem. In order to address the aforementioned problem, in this paper we propose a hashing-based orthogonal deep model to learn accurate and compact multimodal representations. The method can better capture the intra-modality and inter-modality correlations to learn accurate representations. Meanwhile, in order to make the representations compact, the hashing-based model can generate compact hash codes and the proposed orthogonal structure can reduce the redundant information lying in the codes by imposing orthogonal regularizer on the weighting matrices. We also theoretically prove that, in this case, the learned codes are guaranteed to be approximately orthogonal. Moreover, considering the different characteristics of different modalities, effective representations can be attained with different number of layers for different modalities. Comprehensive experiments on three real-world datasets demonstrate a substantial gain of our method on retrieval tasks compared with existing algorithms.
Keyword:
Deep learning
multimodal hashing
similarity search
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
引用论文

引用论文

Human leucocyte antigen and TNFα polymorphism association in microscopic colitis
err2008-04-01
err0
PREAI
errRitva M. Koskela; Tuomo J. Karttunen; Seppo E. Niemelä; Juhani K. Lehtola; Jorma Ilonen; Riitta A. Karttunen
err分享
err收藏
Semantic hashing
err2009-07-01
err939
errOAAI
errSalakhutdinov, Ruslan; Hinton, Geoffrey
err分享
err收藏
An Organometallic Route to Oligonucleotides Containing Phosphoroselenoate
err2002-11-04
err0
PREAI
errGeoffrey A. Holloway; Caroline Pavot; Stephen A. Scaringe; Yi Lu; Thomas B. Rauchfuss
err分享
err收藏
学者 查看更多内容