Return
Reframing Audio Data Annotation as Domain Adaptation Process: A Multi-Indicator Active Learning Framework
DOI:10.1109/TASLPRO.2025.3648795.png)
Abstract
En 中文
Annotated datasets are critical for training high-performance audio models, yet manual annotation remains labor-intensive and costly. While Active Learning (AL) is proposed to annotate samples selectively to maximize model performance, most existing approaches are based on the pre-defined set of categories, called the fixed-set paradigm. This fixed label set cannot effectively address the dynamic annotation needed in the expanding-set paradigm, where new sound classes emerge over time. To bridge this gap, this work reframes audio annotation as a domain adaptation process to unify these two paradigms and resolves them by a general framework. Unlike traditional AL methods that rely on a single indicator, this work proposes a Multi-Indicator Active Learning Framework (MIAF) that integrates multiple indicators, to meet different needs of fixed-set and expanding-set paradigms. The MIAF introduces a novel indicator, Domain Shift Impact (DSI), to minimize the distributional divergence between labeled data and the full dataset. Based on the DSI, the indicator of diversity is introduced to prevent information redundancy in a batch sampling, while incorporating indicators of density and outlierness as paradigm-specific indicators for fixed-set and expanding-set scenarios, respectively. To integrate these indicators, a region-based sample selection strategy is proposed to separate the feature space into several regions depending on density or outlierness, while regions with higher DSI value are prioritized, with one sample selected from each region to ensure diversity. Experiments on four diverse audio datasets demonstrate that MIAF consistently outperforms existing AL methods across different paradigms and dataset sizes. Ablation studies further validate the complementary roles of the proposed indicators and the generalization capacity of MIAF.
Keywords:
Annotations
Active learning
Labeling
Adaptation models
Speech recognition
Training
Event detection
Data models
Uncertainty
Speech enhancement
Data annotation
active learning (AL)
audio tagging
Journal
I
IF:
0
Papers:
151
Citations:
0

