arrow
返回

Speaker Adaptive Text-to-Speech With Timbre-Normalized Vector-Quantized Feature

delete2023-01-01
delete2
PRE
AI
C
Chenpeng Du
Y
Yiwei Guo
陈
陈谐 (Xie Chen)
K
Kai Yu *
DOI:10.1109/TASLP.2023.3308374delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Achieving high fidelity and speaker similarity in text-to-speech speaker adaptation with limited amount of data is a challenging task. Most existing methods only consider adapting to the timbre of the target speakers but fail to capture their speaking styles from little data. In this work, we propose a novel TTS system, TN-VQTTS, which leverages timbre-normalized vector-quantized (TN-VQ) acoustic feature for speaker adaptation with little data. With the TN-VQ feature, speaking style and timbre can be effectively decomposed and controlled by the acoustic model and the vocoder separately of VQTTS. Such decomposition enables us to finely mimic both the two characteristics of the target speaker in adaptation with little data. Specifically, we first reduce the dimensionality of self-supervised VQ acoustic feature via PCA and normalize its timbre with a normalizing flow model. The feature is then quantized with k-means and used as the TN-VQ feature for a multi-speaker VQ-TTS system. Furthermore, we optimize timbre-independent style embeddings of the training speakers jointly with the acoustic model and store them in a lookup table. The embedding table later serves as a selectable codebook or a group of basis for representing the style of unseen speakers. Our experiments on LibriTTS dataset first show that the proposed model architecture for VQ feature achieves better performance in multi-speaker text-to-speech synthesis than several existing methods. We also find that the reconstruction performance and the naturalness are almost unchanged after applying timbre normalization and k-means quantization. Finally, we show that TN-VQTTS achieves better performance on speaker similarity in adaptation than both speaker embedding based adaptation method and fine-tuning based baseline AdaSpeech.
Keyword:
Speech synthesis
speaker adaptation
timbre normalization
vector quantization

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

S
shanghai jiao tong university
学者数:
15.7W
论文数: 11.7W
被引数: 159
引用论文

引用论文

Shapes and Other Things
err2015-08-22
err0
errOAAI
errTerry Knight
err分享
err收藏
Microstructure and physical properties of nano charcoal ash as binder
err2019-04-01
err0
errOAAI
errSiti Nur Amiera Jeffry; Ramadhansyah Putra Jaya; Norhidayah Abdul Hassan; Jahangir Mirza; Mohd Ibrahim Mohd Yusak
err分享
err收藏
Look-Ahead VNF-FG Embedding Framework for Latency-Sensitive Network Services面向低延迟网络服务的提前预览VNF-FG嵌入框架
err2023-09-01
err0
PREAI
errÁkos Recse; Nattakorn Promwongsa; Amin Ebrahimzadeh; Seyedeh Negar Afrasiabi; Carla Mouradian; Wubin Li; Róbert Szabó; Roch H. Glitho
err分享
err收藏
学者 查看更多内容