Return
MSMF-MIL: Multi-Scale Mixed Feature-Based Multiple Instance Learning for Speech Emotion Recognition
DOI:10.1109/TCE.2025.3596056.png)
Abstract
En 中文
Speech emotion in natural utterances is inherently complex and non-uniform, with key emotional cues often confined to brief segments. To address this challenge, we propose a novel multi-scale mixed feature-based framework that leverages Multiple Instance Learning (MIL). In our approach, each speech utterance is transformed into a “bag” containing multiple segments, with each segment treated as an individual instance. Inspired by MIL’s capability to identify key instances within a set, our framework employs CNN-based MIL models at both the frame and utterance-levels, while a ResNet-based model extracts segment-level features. These multi-scale representations are then fused to isolate and emphasize the critical speech segments that express dominant emotions. Experimental evaluations on the Interactive Emotional Dyadic Motion Capture Database (IEMOCAP), the Berlin Database of Emotional Speech (Emo-DB) and the spontaneous URDU-language speech emotion database (URDU) demonstrate our approach’s effectiveness, achieving 67.76% weighted accuracy and 58.39% unweighted accuracy on IEMOCAP, 93.80% weighted accuracy and 91.92% unweighted accuracy on Emo-DB, and 95.75% unweighted accuracy on URDU. These results confirm that our MIL-based, multi-scale feature fusion strategy significantly enhances the robustness and accuracy of speech emotion recognition systems.
Keywords:
Multiple instance learning
speech emotion recognition
multi-scale mixed feature
Journal
IF:
10.9
Papers:
5.1K
Citations:
6.8K

