arrow
返回

A Discriminative Vectorial Framework for Multi-Modal Feature Representation

delete2022-01-01
delete13
delete
OA
AI
高蕾 (Lei Gao) *
L
Ling Guan
DOI:10.1109/TMM.2021.3066118delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Due to the rapid advancements of sensory and computing technology, multi-modal data sources that represent the same pattern or phenomenon have attracted growing attention. As a result, finding means to explore useful information from these multi-modal data sources has quickly become a necessity. In this paper, a discriminative vectorial framework is proposed for multi-modal feature representation in knowledge discovery by employing multi-modal hashing (MH) and discriminative correlation maximization (DCM) analysis. Specifically, the proposed framework is capable of minimizing the semantic similarity among different modalities by MH and exacting intrinsic discriminative representations across multiple data sources by DCM analysis jointly, enabling a novel vectorial framework of multi-modal feature representation. Moreover, the proposed feature representation strategy is analyzed and further optimized based on canonical and non-canonical cases, respectively. Consequently, the generated feature representation leads to effective utilization of the input data sources of high quality, producing improved, sometimes quite impressive, results in various applications. The effectiveness and generality of the proposed framework are demonstrated by utilizing classical features and deep neural network (DNN) based features with applications to image and multimedia analysis and recognition tasks, including data visualization, face recognition, object recognition; cross-modal (text-image) recognition and audio emotion recognition. Experimental results show that the proposed solutions are superior to state-of-the-art statistical machine learning (SML) and DNN algorithms.
Keyword:
Semantics
Correlation
Task analysis
Emotion recognition
Visualization
Transforms
Image recognition
Audio emotion recognition
cross-modal analysis
discriminative correlation maximization
image analysis and recognition
knowledge discovery
multi-modal feature representation
multi-modal hashing
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

T
Toronto Metropolitan University
学者数:
6.0K
论文数: 7.0K
被引数: 6.4K
引用论文

引用论文

SkeletonNet: A Hybrid Network With a Skeleton-Embedding Process for Multi-View Image Representation Learning
err2019-11-01
err40
PREAI
errYang, Shijie; Li, Liang; Wang, Shuhui; Zhang, Weigang; Huang, Qingming; Tian, Qi
err分享
err收藏
Canonical Correlation Analysis With L2,1-Norm for Multiview Data Representation
err2020-11-01
err27
PREAI
errXu, Meixiang; Zhu, Zhenfeng; Zhang, Xingxing; Zhao, Yao; Li, Xuelong
err分享
err收藏
The 1965 Eruption of Taal Volcano
err1966-02-25
err0
PREAI
errJames G. Moore; Kazuaki Nakamura; Arturo Alcaraz
err分享
err收藏
err分享
err收藏
Combination of Silk Fibroin with Acid and with Base
err1941-01-01
err0
PREAI
errLeland F. Gleysteen; Milton Harris
err分享
err收藏
Special Section on Music Data Mining
err2014-08-01
err4
errOAAI
errLi, Tao; Ogihara, Mitsunori; Tzanetakis, George
err分享
err收藏
学者 查看更多内容