返回
Adaptive data augmentation for mandarin automatic speech recognition
DOI:10.1007/s10489-024-05381-6.png)
摘要
En 中文
Audio data augmentation is widely adopted in automatic speech recognition (ASR) to alleviate the overfitting problem. However, noise-based data augmentation converts an over-fitting problem into an under-fitting problem which increases the training time severely. With noise-based data augmentation, informative features are not be persisted during the generating process and generated audio clips would become noise data for the acoustic model. To face the challenge, we propose an Adaptive audio Data Augmentation method called ADA with deep clustering. The proposed ADA could automatically select the most informative augmented sample for each generation. Moreover, two sample selection strategies called RM and RS are proposed. The proposed RM removes samples whose embedding are far away from the cluster center, while the proposed RS maintains the diversity of augmentation samples by sampling in each cluster. Experiments on Aishell-1 demonstrate that the proposed ADA method could improve the data efficiency of end-to-end ASR model in both CNN-based and Transformer-based networks. The proposed ADA obtains an 11.28% and 5.95% relative improvement on SS-CNN and LS-CNN, and a 4.35% improvement on S-Transformer compared with the state-of-the-art audio data augmentation method. Meanwhile, the proposed ADA method decreases the demand of augmented samples by 2.7 times in SS-CNN, LS-CNN and S-Transformer. The qualitative and quantitative analysis proves the effectiveness and efficiency of the proposed ADA method.
Keyword:
Adaptive data augmentation
Data efficiency
Deep clustering
Speech recognition
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
暂无机构信息
引用论文
Improving deep speech denoising by Noisy2Noisy signal mapping利用noisy2noisy2信号映射改进深度语音去噪
APPLIED ACOUSTICS
IF3.6
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units休伯特: 通过隐藏单元的掩蔽预测进行自我监督的语音表示学习

