Return
A Self-Improving Ensemble Learning Framework for Human Non-Speech Vocalization Classification
DOI:10.1109/ACCESS.2025.3639444.png)
Abstract
En 中文
Human non-speech vocalization recognition and classification play a crucial role in emotional and health analysis, but they face significant challenges owing to data scarcity and noisy labels. Although ignoring noisy labels can improve model performance, it also reduces dataset size. To maximize utilization of all available audio data, we propose a self-improving, model-agnostic method based on an ensemble teacher-student framework. This approach leverages semi-supervised ensemble learning techniques to effectively guide student models in learning more features, while utilizing enhanced pseudo-labels derived from ensemble teacher models to significantly reduce the adverse effects of label noise and substantially enhance model performance. Additionally, integrating this method with data augmentation, weight averaging, transfer learning, and other complementary training strategies can further boost accuracy and overall effectiveness. Extensive experimental results on the VocalSound and other datasets confirm that the proposed method not only effectively mitigates label noise in audio datasets—leading to noticeable improvements in accuracy and stability—but also ensures strong generalization across various model architectures. The datasets and codes are available at https://github.com/helensnn/ELVS
Keywords:
Ensemble learning
human non-speech vocalization
pseudo-labels
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.6
Papers:
9.8W
Citations:
29.4W

