Return
Universal Source Separation With Weakly Labelled Data for Computational Auditory Scene Analysis
DOI:10.1109/TASLPRO.2025.3636294.png)
Abstract
En 中文
Computational Auditory Scene Analysis (CASA) aims to detect what sounds appear and when they occur in an audio recording, as well as to separate the audio into waveforms corresponding to the identified sound classes. The CASA problem presents several challenges. First, previous source separation systems require training on clean data, which is often scarce. Second, existing conditional source separation systems are designed to separate only a limited number of sound classes and cannot scale to separate hundreds of classes. Third, prior works require users to specify the sounds to be separated and cannot automatically detect which sound classes to separate. Fourth, there is a limited research on building hierarchical source separation systems. The contributions of this work are as follows. First, we propose training universal source separation (USS) systems on large-scale weakly labeled and unlabelled datasets, rather than relying solely on clean data. Second, we develop a large-scale USS system capable of separating up to 527 sound classes, with the potential to scale to unlimited number of classes. Third, we introduce a method to automatically detect and separate sound classes without user specification. Fourth, we present a hierarchical USS system that can separate sound classes at various hierarchical levels. We train the USS systems on AudioSet and evaluate their performance on speech, audio, and music datasets, demonstrating the effectiveness of our approach.
Keywords:
Source separation
Training
Data mining
Time-domain analysis
Neural networks
Time-frequency analysis
Tagging
Fourier transforms
Convolutional neural networks
Transformers
Universal source separation
computational auditory scene analysis
weakly labelled data
hierarchical source separation
Journal
I
IF:
0
Papers:
151
Citations:
0

