返回
Evaluating deep learning architectures for Speech Emotion Recognition
DOI:10.1016/j.neunet.2017.02.013.png)
摘要
En 中文
Speech Emotion Recognition (SER) can be regarded as a static or dynamic classification problem, which makes SER an excellent test bed for investigating and comparing various deep learning architectures. We describe a frame-based formulation to SER that relies on minimal speech processing and end-to-end deep learning to model intra-utterance dynamics. We use the proposed SER system to empirically explore feed-forward and recurrent neural network architectures and their variants. Experiments conducted illuminate the advantages and limitations of these architectures in paralinguistic speech recognition and emotion recognition in particular. As a result of our exploration, we report state-of-the-art results on the IEMOCAP database for speaker-independent SER and present quantitative and qualitative assessments of the models' performances. (C) 2017 Elsevier Ltd. All rights reserved.
Keyword:
Affective computing
Deep learning
Emotion recognition
Neural networks
Speech recognition
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.3
论文数:
8.2K
被引数:
3.0W
机构
暂无机构信息
引用论文
Survey on speech emotion recognition: Features, classification schemes, and databases
PATTERN RECOGNITION
IF7.6
Two-dimensional bricklayer arrangements of tolans using halogen bonding interactions使用卤素键相互作用的tolans的二维瓦工层布置

