arrow
返回

Rhythm Modeling for Voice Conversion

delete2023-01-01
delete0
delete
OA
AI
B
Benjamin van Niekerk *
M
Marc‐André Carbonneau
H
Herman Kamper
DOI:10.1109/LSP.2023.3313515delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we introduce Urhythmic-an unsupervised method for rhythm conversion that does not require parallel data or text transcriptions. Using self-supervised representations, we first divide source audio into segments approximating sonorants, obstruents, and silences. Then we model rhythm by estimating speaking rate or the duration distribution of each segment type. Finally, we match the target speaking rate or rhythm by time-stretching the speech segments. Experiments show that Urhythmic outperforms existing unsupervised methods in terms of quality and prosody.
Keyword:
Voice conversion
rhythm conversion
speaking rate estimation

期刊

IEEE Signal Processing Magazine 封面图
IEEE Signal Processing Magazine
IF:
9.6
论文数:
1.1W
被引数:
1.7W

机构

暂无机构信息
引用论文

引用论文

Self-Supervised Speech Representation Learning: A Review自我监督的语音表示学习: 综述
err2022-10-01
err92
errOAAI
errMohamed, Abdelrahman; Lee, Hung-yi; Borgholt, Lasse; Havtorn, Jakob D.; Edin, Joakim; Igel, Christian; Kirchhoff, Katrin; Li, Shang-Wen; Livescu, Karen; Maaloe, Lars; Sainath, Tara N.; Watanabe, Shinji
err分享
err收藏
err分享
err收藏
学者 查看更多内容