arrow
返回

Robust processing techniques for voice conversion

delete2006-10-01
delete41
PRE
AI
O
Oytun Türk *
L
Levent M. Arslan
DOI:10.1016/j.csl.2005.06.001delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Differences in speaker characteristics, recording conditions, and signal processing algorithms affect. output quality in voice conversion systems. This study focuses on formulating robust techniques for a codebook mapping based voice conversion algorithm. Three different methods are used to improve voice conversion performance: confidence measures, pre-emphasis, and, spectral equalization. Analysis is performed for each method and the implementation details are discussed. The first method employs confidence measures in the training stage to eliminate problematic pairs of source and target speech units that might result from possible misalignments, speaking style differences or pronunciation variations. Four confidence measures are developed based on the spectral distance, fundamental frequency (f(0)) distance, energy distance, and duration distance between the source and target speech units. The second method focuses on the importance of pre-emphasis in line-spectral frequency (LSF) based vocal tract modeling and transformmation. The last method, spectral equalization, is aimed at reducing the differences in the source and target long-term spectra when the source and target recording conditions are significantly different. The voice conversion algorithm that employs the proposed techniques, is compared with the baseline voice conversion algorithm with objective tests as well as three subjective listening tests. First, similarity to the target voice is evaluated in a subjective listening test and it is shown that the proposed algorithm improves similarity to the target voice by 23.0%. An ABX test is performed and the proposed algorithm is preferred over the baseline algorithm by 76.4%. In the third test, the two algorithms are compared in terms of the subjective quality. of the voice conversion output. The proposed algorithm improves the subjective output quality by 46.8% in terms of mean opinion score (MOS). (c) 2005 Elsevier Ltd. All rights reserved.
Keyword:
TRANSFORMATION
INDIVIDUALITY
RECOGNITION
PREDICTION
TUTORIAL
QUALITY
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

C
Computer Speech and Language
IF:
3.4
论文数:
1.5K
被引数:
2.6K

机构

暂无机构信息
引用论文

引用论文

Coupled membrane transporters reduce noise
err2020-01-27
err0
errOAAI
errLuca Cardelli; Luca Laurenti; Attila Csikasz-Nagy
err分享
err收藏
Sulphasomizole (5-p-Aminobenzenesulphonamido-3-Methylisothiazole): A New Antibacterial Sulphonamide
err1960-04-01
err0
PREAI
errA. ADAMS; W. A. FREEMAN; A. HOLLAND; D. HOSSACK; J. INGLIS; J. PARKINSON; H. W. READING; K. RIVETT; R. SLACK; R. SUTHERLAND; R. WIEN
err分享
err收藏
学者 查看更多内容