返回
Rhythm Modeling for Voice Conversion
DOI:10.1109/LSP.2023.3313515.png)
摘要
En 中文
Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we introduce Urhythmic-an unsupervised method for rhythm conversion that does not require parallel data or text transcriptions. Using self-supervised representations, we first divide source audio into segments approximating sonorants, obstruents, and silences. Then we model rhythm by estimating speaking rate or the duration distribution of each segment type. Finally, we match the target speaking rate or rhythm by time-stretching the speech segments. Experiments show that Urhythmic outperforms existing unsupervised methods in terms of quality and prosody.
Keyword:
Voice conversion
rhythm conversion
speaking rate estimation
期刊
IF:
9.6
论文数:
1.1W
被引数:
1.7W
机构
暂无机构信息
引用论文
Loss of heterozygosity in yeast can occur by ultraviolet irradiation during the S phase of the cell cycle酵母中的杂合性丢失可以通过细胞周期的S期中的紫外线照射发生。

