返回
Spontaneous Speech Emotion Recognition Using Multiscale Deep Convolutional LSTM
DOI:10.1109/TAFFC.2019.2947464.png)
摘要
En 中文
Recently, emotion recognition in real sceneries such as in the wild has attracted extensive attention in affective computing, because existing spontaneous emotions in real sceneries are more challenging and difficult to identify than other emotions. Motivated by the diverse effects of different lengths of audio spectrograms on emotion identification, this paper proposes a multiscale deep convolutional long short-term memory (LSTM) framework for spontaneous speech emotion recognition. Initially, a deep convolutional neural network (CNN) model is used to learn deep segment-level features on the basis of the created image-like three channels of spectrograms. Then, a deep LSTM model is adopted on the basis of the learned segment-level CNN features to capture the temporal dependency among all divided segments in an utterance for utterance-level emotion recognition. Finally, different emotion recognition results, obtained by combining CNN with LSTM at multiple lengths of segment-level spectrograms, are integrated by using a score-level fusion strategy. Experimental results on two challenging spontaneous emotional datasets, i.e., the AFEW5.0 and BAUM-1s databases, demonstrate the promising performance of the proposed method, outperforming state-of-the-art methods.
Keyword:
Image segmentation
Spectrogram
Feature extraction
Emotion recognition
Speech recognition
Neural networks
Acoustics
Speech emotion recognition
convolutional neural networks
LSTM
multiscale
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.8
论文数:
1.4K
被引数:
9.1K
机构
引用论文
Audio-visual emotion fusion (AVEF): A deep efficient weighted approach视听情感融合 (AVEF): 一种深度高效加权方法
INFORMATION FUSION
IF15.5

