arrow
Return

Adaptive data augmentation for mandarin automatic speech recognition

delete2024-04-24
delete2
PRE
AI
K
Kai Ding *
R
Ruixuan Li *
X
Xu, Yuelin
X
Xingyue Du
邓斌 cover
邓斌 (Bin Deng)
DOI:10.1007/s10489-024-05381-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Audio data augmentation is widely adopted in automatic speech recognition (ASR) to alleviate the overfitting problem. However, noise-based data augmentation converts an over-fitting problem into an under-fitting problem which increases the training time severely. With noise-based data augmentation, informative features are not be persisted during the generating process and generated audio clips would become noise data for the acoustic model. To face the challenge, we propose an Adaptive audio Data Augmentation method called ADA with deep clustering. The proposed ADA could automatically select the most informative augmented sample for each generation. Moreover, two sample selection strategies called RM and RS are proposed. The proposed RM removes samples whose embedding are far away from the cluster center, while the proposed RS maintains the diversity of augmentation samples by sampling in each cluster. Experiments on Aishell-1 demonstrate that the proposed ADA method could improve the data efficiency of end-to-end ASR model in both CNN-based and Transformer-based networks. The proposed ADA obtains an 11.28% and 5.95% relative improvement on SS-CNN and LS-CNN, and a 4.35% improvement on S-Transformer compared with the state-of-the-art audio data augmentation method. Meanwhile, the proposed ADA method decreases the demand of augmented samples by 2.7 times in SS-CNN, LS-CNN and S-Transformer. The qualitative and quantitative analysis proves the effectiveness and efficiency of the proposed ADA method.
Keywords:
Adaptive data augmentation
Data efficiency
Deep clustering
Speech recognition

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.6K
Citations:
1.7W

Organization

No organization information available
Cited Papers

Cited Papers

Improving deep speech denoising by Noisy2Noisy signal mapping
err2021-01-01
err32
errOAAI
errAlamdari, N.; Azarang, A.; Kehtarnavaz, N.
errShare
errSave
Deep CNN-based local dimming technologys
err2021-05-12
err12
PREAI
errZhang, Tao; Wang, Hao; Du, Wenli; Li, Meng
errShare
errSave
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
err2021-01-01
err1.1K
errOAAI
errHsu, Wei-Ning; Bolte, Benjamin; Tsai, Yao-Hung Hubert; Lakhotia, Kushal; Salakhutdinov, Ruslan; Mohamed, Abdelrahman
errShare
errSave
ASTT: acoustic spatial-temporal transformer for short utterance speaker recognition
err2023-03-04
err7
PREAI
errWu, Xing; Li, Ruixuan; Deng, Bin; Zhao, Ming; Du, Xingyue; Wang, Jianjia; Ding, Kai
errShare
errSave
errShare
errSave
The Anchor Borrowers Programme and youth rice farmers in Northern Nigeria
err2020-08-27
err0
PREAI
errOpeyemi Olanrewaju; Romanus Osabohien; James Fasakin
errShare
errSave
errShare
errSave
Perceived Problems of Secondary School Teachers1
err2014-12-07
err0
PREAI
errDonald R. Cruickshank; John J. Kennedy; Betty Myers
errShare
errSave
An Efficient Data Augmentation Network for Out-of-Distribution Image Detection
err2021-01-01
err9
errOAAI
errLin, Cheng-Hung; Lin, Cheng-Shian; Chou, Po-Yung; Hsu, Chen-Chien
errShare
errSave
researcher View more