arrow
Return

Machine Speech Chain

delete2020-01-01
delete25
delete
OA
AI
T
Tjandra, Andros *
S
Sakriani Sakti
S
Satoshi Nakamura
DOI:10.1109/TASLP.2020.2977776delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Despite the close relationship between speech perception and production, research in automatic speech recognition (ASR) and text-to-speech synthesis (TTS) has progressed more or less independently without exerting much mutual influence. In human communication, on the other hand, a closed-loop speech chain mechanism with auditory feedback from the speaker's mouth to her ear is crucial. In this paper, we take a step further and develop a closed-loop machine speech chain model based on deep learning. The sequence-to-sequence model in closed-loop architecture allows us to train our model on the concatenation of both labeled and unlabeled data. While ASR transcribes the unlabeled speech features, TTS attempts to reconstruct the original speech waveform based on the text from ASR. In the opposite direction, ASR also attempts to reconstruct the original text transcription given the synthesized speech. To the best of our knowledge, this is the first deep learning framework that integrates human speech perception and production behaviors. Our experimental results show that the proposed approach significantly improved performance over that from separate systems that were only trained with labeled data.
Keywords:
Speech processing
Data models
Machine learning
Task analysis
Training
Hidden Markov models
Speech chain
ASR
TTS
deep learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

N
nara institute of science & technology
Scholars:
4.1K
Papers: 3.1K
Citations: 7