返回
A deep interpretable representation learning method for speech emotion recognition
DOI:10.1016/j.ipm.2023.103501.png)
摘要
En 中文
This paper focuses on the active interpretability for deep learning-based speech emotion recognition (SER). To achieve this, we propose an explicit feature constrained model, the interpretable group convolutional neural network (IG-CNN) model. In the proposed model, we first introduce the interpretability constraint to learn human-understandable interpretable representations. The emotion prediction decision can be active interpreted via the model coefficients. To acquire more representations beyond interpretable ones, and ensure they are useful for SER, we then design the uncorrelation constraint between interpretable and autonomous representations and introduce group CNN structure. We test the model on IEMOCAP, RAVDESS, eNTERFACE'05, and CREMA-D datasets. Experimental results show that our model outperforms all the baselines. In addition, the proposed model can also learn the patterns of human perception of speech emotion and provide explanation for the recognition results.
Keyword:
Speech signals
Emotion recognition
Affective computing
Interpretability
Representation learning
Deep learning
期刊
I
IF:
6.9
论文数:
5.2K
被引数:
1.4W
机构
引用论文
Gender Differences in Implicit and Explicit Processing of Emotional Facial Expressions as Revealed by Event-Related Theta Synchronization
EMOTION
IF3.6
Spontaneous Speech Emotion Recognition Using Multiscale Deep Convolutional LSTM基于多尺度深度卷积LSTM的自发语音情感识别

