返回
Iterative Feature Normalization Scheme for Automatic Emotion Detection from Speech
DOI:10.1109/T-AFFC.2013.26.png)
摘要
En 中文
The externalization of emotion is intrinsically speaker-dependent. A robust emotion recognition system should be able to compensate for these differences across speakers. A natural approach is to normalize the features before training the classifiers. However, the normalization scheme should not affect the acoustic differences between emotional classes. This study presents the iterative feature normalization (IFN) framework, which is an unsupervised front-end, especially designed for emotion detection. The IFN approach aims to reduce the acoustic differences, between the neutral speech across speakers, while preserving the inter-emotional variability in expressive speech. This goal is achieved by iteratively detecting neutral speech for each speaker, and using this subset to estimate the feature normalization parameters. Then, an affine transformation is applied to both neutral and emotional speech. This process is repeated till the results from the emotion detection system are consistent between consecutive iterations. The IFN approach is exhaustively evaluated using the IEMOCAP database and a data set obtained under free uncontrolled recording conditions with different evaluation configurations. The results show that the systems trained with the IFN approach achieve better performance than systems trained either without normalization or with global normalization.
Keyword:
Emotion recognition
speaker normalization
emotion
features normalization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.8
论文数:
1.4K
被引数:
9.1K
机构
引用论文
Micelle-like particles formed by carboxylic acid-terminated polystyrene and poly(4-vinyl pyridine) in chloroform/methanol mixed solution
Polymer
IF0
Effects of Thiosulfate on Susceptibility of Type 316 Stainless Steel to Stress Corrosion Cracking in 3.5% Aqueous Sodium Chloride
CORROSION
IF0

