arrow
返回

Joint Speech-Text Embeddings for Multitask Speech Processing

delete2024-01-01
delete0
delete
OA
AI
M
Michael Gian Gonzales *
P
Peter Corcoran
N
Naomi Harte
M
Michael Schukat
DOI:10.1109/ACCESS.2024.3473743delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Devices that use speech as the communication medium between human and computer have been emerging for the past few years. The technologies behind this interface are called Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The two are distinct fields in speech signal processing that have independently made great strides in recent years. This paper proposes an architecture that takes advantage of the two modalities present in ASR and TTS, speech and text, while simultaneously training three tasks, adding speaker recognition to the underlying ASR and TTS tasks. This architecture not only reduces the memory footprint required to run all tasks, but also has performance comparable to single-task models. The dataset used to train and evaluate the model is the CSTR VCTK Corpus. Results show a 97.64% accuracy in the speaker recognition task, word and character error rates of 18.18% and 7.95% for the ASR task, a mel cepstral distortion of 4.31 and two predicted MOS of 2.98 and 3.28 for the TTS task. While voice conversion is not part of the training tasks, the architecture is capable of doing this and was evaluated to have 5.22, 2.98, and 2.73 for mel cepstral distortion and predicted MOS, respectively.
Keyword:
Speaker recognition
Spectrogram
Decoding
Training
Feature extraction
Analytical models
Speech coding
Propagation losses
Automatic speech recognition
Speech to text
Text to speech
Speech processing
joint speech-text
text-to-speech
speaker recognition
speech processing
voice conversion

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

T
Trinity College Dublin
学者数:
2.4W
论文数: 1.9W
被引数: 2.7W
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Wave Propagation and Diffraction
err2018-01-01
err0
PREAI
errIgor T. Selezov; Yuriy G. Kryvonos; Ivan S. Gandzha
err分享
err收藏
An information set-based robust text-independent speaker authentication
err2019-08-14
err4
PREAI
errMedikonda, Jeevan; Bhardwaj, Saurabh; Madasu, Hanmandlu
err分享
err收藏
Machine Speech Chain
err2020-01-01
err25
errOAAI
errTandra, Andros; Sakti, Sakriani; Nakamura, Satoshi
err分享
err收藏
学者 查看更多内容