arrow
返回

Self-Supervised Pre-Training for Attention-Based Encoder-Decoder ASR Model

delete2022-01-01
delete3
PRE
AI
C
Changfeng Gao
G
Gaofeng Cheng *
L
Li Ta
P
Pengyuan Zhang
Y
Yonghong Yan
DOI:10.1109/TASLP.2022.3171967delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
End-to-end (E2E) models, including the attention-based encoder-decoder (AED) models, have achieved promising performance on the automatic speech recognition (ASR) task. However, the supervised training process of the E2E model needs a large amount of speech-text paired data. In contrast, self-supervised pre-training can pre-train the model on the unlabeled data and then fine-tune it on the limited labeled data to realize better performance. Most of the previous self-supervised pre-training methods focus on learning hidden representations from speech but ignore how to utilize the unpaired text. As a result, previous works often pre-train an acoustic encoder and then fine-tune it as a classification based ASR model, such as Connectionist Temporal Classification (CTC) based model, rather than an AED model. In this paper, we propose a self-supervised pre-training method for the AED model (SP-AED). The SP-AED method contains acoustic pre-training for the encoder, linguistic pre-training for the decoder, and an adaptive combination fine-tuning for the whole system. We first design a linguistic pre-training method for decoder by utilizing the text-only data. The decoder will be pre-trained as a noise-condition language model to learn the prior distribution of the text. Then, we pre-train the AED encoder with the wav2vec2.0 method with some modifications. Finally, we combine the pre-trained encoder and decoder and fine-tune them on the limited labeled data. We design an adaptive combination method during fine-tuning by modifying the decoder's input and output to prevent catastrophic forgetting. Experiments prove that compared with the random initialized models, the SP-AED pre-trained models can realize up to 17% relative improvement. And with similar model size or computational cost, we can get comparable results to other classification-based models on both English and Chinese corpus.
Keyword:
Decoding
Acoustics
Data models
Training
Linguistics
Feature extraction
Computational modeling
End-to-end
self-supervised pre-training
speech recognition

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Ribosomes in platelets protect the messenger
err2017-04-27
err0
errOAAI
errJesse W. Rowley; Andrew S. Weyrich
err分享
err收藏
Impurity effects on ionic-liquid-based supercapacitors
err2016-12-27
err0
errOAAI
errKun Liu; Cheng Lian; Douglas Henderson; Jianzhong Wu
err分享
err收藏
BCI, an inhibitor of the DUSP1 and DUSP6 dual specificity phosphatases, enhances P2X7 receptor expression in neuroblastoma cells
err2022-12-15
err0
errOAAI
errMaría Benito-León; Juan Carlos Gil-Redondo; Raquel Perez-Sen; Esmerilda G. Delicado; Felipe Ortega; Rosa Gomez-Villafuertes
err分享
err收藏
MACULAR ATROPHY FINDINGS BY OPTICAL COHERENCE TOMOGRAPHY ANGIOGRAPHY COMPARED WITH FUNDUS AUTOFLUORESCENCE IN TREATED EXUDATIVE AGE-RELATED MACULAR DEGENERATION
err2019-02-01
err0
errOAAI
errYukari Takasago; Chieko Shiragami; Mamoru Kobayashi; Rie Osaka; Aoi Ono; Ayana Yamashita; Akitaka Tsujikawa; Kazuyuki Hirooka
err分享
err收藏
YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
Hybrid CTC/Attention Architecture for End-to-End Speech Recognition
err2017-12-01
err512
PREAI
errWatanabe, Shinji; Hori, Takaaki; Kim, Suyoun; Hershey, John R.; Hayashi, Tomoki
err分享
err收藏
err分享
err收藏
学者 查看更多内容