arrow
返回

Adaptive data augmentation for mandarin automatic speech recognition

delete2024-04-24
delete2
PRE
AI
K
Kai Ding *
R
Ruixuan Li *
X
Xu, Yuelin
X
Xingyue Du
邓斌 封面图
邓斌 (Bin Deng)
DOI:10.1007/s10489-024-05381-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Audio data augmentation is widely adopted in automatic speech recognition (ASR) to alleviate the overfitting problem. However, noise-based data augmentation converts an over-fitting problem into an under-fitting problem which increases the training time severely. With noise-based data augmentation, informative features are not be persisted during the generating process and generated audio clips would become noise data for the acoustic model. To face the challenge, we propose an Adaptive audio Data Augmentation method called ADA with deep clustering. The proposed ADA could automatically select the most informative augmented sample for each generation. Moreover, two sample selection strategies called RM and RS are proposed. The proposed RM removes samples whose embedding are far away from the cluster center, while the proposed RS maintains the diversity of augmentation samples by sampling in each cluster. Experiments on Aishell-1 demonstrate that the proposed ADA method could improve the data efficiency of end-to-end ASR model in both CNN-based and Transformer-based networks. The proposed ADA obtains an 11.28% and 5.95% relative improvement on SS-CNN and LS-CNN, and a 4.35% improvement on S-Transformer compared with the state-of-the-art audio data augmentation method. Meanwhile, the proposed ADA method decreases the demand of augmented samples by 2.7 times in SS-CNN, LS-CNN and S-Transformer. The qualitative and quantitative analysis proves the effectiveness and efficiency of the proposed ADA method.
Keyword:
Adaptive data augmentation
Data efficiency
Deep clustering
Speech recognition

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

暂无机构信息
引用论文

引用论文

Deep CNN-based local dimming technologys
err2021-05-12
err12
PREAI
errZhang, Tao; Wang, Hao; Du, Wenli; Li, Meng
err分享
err收藏
ASTT: acoustic spatial-temporal transformer for short utterance speaker recognition
err2023-03-04
err7
PREAI
errWu, Xing; Li, Ruixuan; Deng, Bin; Zhao, Ming; Du, Xingyue; Wang, Jianjia; Ding, Kai
err分享
err收藏
err分享
err收藏
The Anchor Borrowers Programme and youth rice farmers in Northern Nigeria
err2020-08-27
err0
PREAI
errOpeyemi Olanrewaju; Romanus Osabohien; James Fasakin
err分享
err收藏
err分享
err收藏
Perceived Problems of Secondary School Teachers1
err2014-12-07
err0
PREAI
errDonald R. Cruickshank; John J. Kennedy; Betty Myers
err分享
err收藏
An Efficient Data Augmentation Network for Out-of-Distribution Image Detection
err2021-01-01
err9
errOAAI
errLin, Cheng-Hung; Lin, Cheng-Shian; Chou, Po-Yung; Hsu, Chen-Chien
err分享
err收藏
学者 查看更多内容