返回
A deep semantic framework for multimodal representation learning
DOI:10.1007/s11042-016-3380-8.png)
摘要
En 中文
Multimodal representation learning has gained increasing importance in various real-world multimedia applications. Most previous approaches focused on exploring inter-modal correlation by learning a common or intermediate space in a conventional way, e.g. Canonical Correlation Analysis (CCA). These works neglected the exploration of fusing multiple modalities at higher semantic level. In this paper, inspired by the success of deep networks in multimedia computing, we propose a novel unified deep neural framework for multimodal representation learning. To capture the high-level semantic correlations across modalities, we adopted deep learning feature as image representation and topic feature as text representation respectively. In joint model learning, a 5-layer neural network is designed and enforced with a supervised pre-training in the first 3 layers for intra-modal regularization. The extensive experiments on benchmark Wikipedia and MIR Flickr 25K datasets show that our approach achieves state-of-the-art results compare to both shallow and deep models in multimodal and cross-modal retrieval.
Keyword:
Multimodal representation
Deep neural networks
Semantic feature
Cross-modal retrieval
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3
论文数:
2.0W
被引数:
3.2W
机构
引用论文
Dynamic Changes in Myofibroblasts Affect the Carcinogenesis and Prognosis of Bladder Cancer Associated With Tumor Microenvironment Remodeling肌成纤维细胞的动态变化与肿瘤微环境重塑对膀胱癌发生及预后的影响

