Return
Multi-task learning for audio scene source counting and analysis
DOI:10.1016/j.mlwa.2025.100785.png)
Abstract
En 中文
Audio source counting is a fundamental task of audio scene analysis related to other audio tasks such as speaker diarization and sound event detection. It is also a relatively unexplored audio task that presents a complex challenge. In particular, source counting performance is poor when the source count range is large, limiting its potential applications. This paper presents a novel approach to improve upon audio source counting through multi-task learning. We present a first of its kind empirical study on the hierarchical nature of audio source counting, introducing the coarse source counting task and a hierarchical multi-task learning framework, in order to better understand and investigate the audio source counting task through several case study scenarios. We perform multi-task learning with a ResNet architecture and demonstrate improvements to audio source counting accuracy by up to a 6% increase from the previous best result on the SARdBScene dataset. We also perform multi-task learning of audio source counting and acoustic scene classification as a step forward for robust audio scene analysis. These experimental results show improvements of up to 6% in source counting accuracy over state-of-the-art baselines, particularly in high source count scenarios. Our findings highlight that multi-task learning not only enhances accuracy, but also improves efficiency by replacing multiple task-specific models with a single robust network.
Keywords:
Acoustic scene classification
Audio source counting
Audio scene analysis
Multi-task learning
ResNet
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
M
IF:
4.9
Papers:
135
Citations:
0

