arrow
返回

Cross-Utterance Conditioned VAE for Speech Generation

delete2024-01-01
delete0
PRE
AI
Y
Yang Li
C
Cheng Yu
G
Guangzhi Sun
W
Weiqin Zu
Z
Zheng Tian
Y
Ying Wen
W
Wei Pan
C
Chao Zhang
J
Jun Wang
杨阳 封面图
杨阳 (Yang Yang)
F
Fanglei Sun *
DOI:10.1109/TASLP.2024.3453598delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Speech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In response, we present the Cross-Utterance Conditioned Variational Autoencoder speech synthesis (CUC-VAE S2) framework to enhance prosody and ensure natural speech generation. This framework leverages the powerful representational capabilities of pre-trained language models and the re-expression abilities of variational autoencoders (VAEs). The core component of the CUC-VAE S2 framework is the cross-utterance CVAE, which extracts acoustic, speaker, and textual features from surrounding sentences to generate context-sensitive prosodic features, more accurately emulating human prosody generation. We further propose two practical algorithms tailored for distinct speech synthesis applications: CUC-VAE TTS for text-to-speech and CUC-VAE SE for speech editing. The CUC-VAE TTS is a direct application of the framework, designed to generate audio with contextual prosody derived from surrounding texts. On the other hand, the CUC-VAE SE algorithm leverages real mel spectrogram sampling conditioned on contextual information, producing audio that closely mirrors real sound and thereby facilitating flexible speech editing based on text such as deletion, insertion, and replacement. Experimental results on the LibriTTS datasets demonstrate that our proposed models significantly enhance speech synthesis and editing, producing more natural and expressive speech.
Keyword:
Multimedia systems
Neural networks
Natural languages
Production
Speech enhancement
Feature extraction
Acoustics
Text to speech
Mirrors
Spectrogram
Pre-trained language model
speech editing
speech synthesis
TTS
variational autoencoder

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
S
shanghai jiao tong university
学者数:
15.7W
论文数: 11.7W
被引数: 159
U
University College London
学者数:
7.9W
论文数: 6.2W
被引数: 15.7W
U
University of Cambridge
学者数:
7.7W
论文数: 7.1W
被引数: 13.7W
U
university of london
学者数:
21.5W
论文数: 19.7W
被引数: 305
S
ShanghaiTech University
学者数:
9.7K
论文数: 5.9K
被引数: 1.6W
U
University of Manchester
学者数:
5.7W
论文数: 5.3W
被引数: 7.4W
学者 查看更多机构
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err2000-01-01
err0
PREAI
errH. S. Kalsi; M. Dutta; S. K. Sharda; G. K. Padam; N. K. Arora; B. K. Das
err分享
err收藏
VoCo: Text-based Insertion and Replacement in Audio NarrationVoCo: 音频旁白中基于文本的插入和替换
err2017-07-20
err39
PREAI
errJin, Zeyu; Mysore, Gautham J.; Diverdi, Stephen; Lu, Jingwan; Finkelstein, Adam
err分享
err收藏
err分享
err收藏
Evolving Minimally Invasive Techniques for Tear Trough Enhancement
err2015-07-01
err0
PREAI
errRobert H. Hill; Craig N. Czyz; Srinivas Kandapalli; Sandy X. Zhang-Nunes; Kenneth V. Cahill; Allan E. Wulc; Jill A. Foster
err分享
err收藏
Efficacy and Toxicity of Vincristine and CYP3A5 Genetic Polymorphism in Rhabdomyosarcoma Pediatric Egyptian Patients
err2024-04-01
err0
errOAAI
errNorhan Shalaby; Hala Zaki; Osama Badary; Sherif Kamal; Mohamed Nagy; Dalia Makhlouf; Amr Elnashar; Enas Elnadi; Sameh Abdelshafi; Sherif Abouelnaga; Mona Saber
err分享
err收藏
学者 查看更多内容