arrow
返回

A noise-robust voice conversion method with controllable background sounds

delete2024-02-29
delete0
delete
OA
AI
L
Lele Chen
X
Xiongwei Zhang
Y
Yihao Li
M
Meng Sun *
W
Weiwei Chen
DOI:10.1007/s40747-024-01375-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Background noises are usually treated as redundant or even harmful to voice conversion. Therefore, when converting noisy speech, a pretrained module of speech separation is usually deployed to estimate clean speech prior to the conversion. However, this can lead to speech distortion due to the mismatch between the separation module and the conversion one. In this paper, a noise-robust voice conversion model is proposed, where a user can choose to retain or to remove the background sounds freely. Firstly, a speech separation module with a dual-decoder structure is proposed, where two decoders decode the denoised speech and the background sounds, respectively. A bridge module is used to capture the interactions between the denoised speech and the background sounds in parallel layers through information exchanging. Subsequently, a voice conversion module with multiple encoders to convert the estimated clean speech from the speech separation model. Finally, the speech separation and voice conversion module are jointly trained using a loss function combining cycle loss and mutual information loss, aiming to improve the decoupling efficacy among speech contents, pitch, and speaker identity. Experimental results show that the proposed model obtains significant improvements in both subjective and objective evaluation metrics compared with the existing baselines. The speech naturalness and speaker similarity of the converted speech are 3.47 and 3.43, respectively.
Keyword:
Noise-robust voice conversion
Dual-decoder structure
Bridge module
Cycle loss
Speech disentanglement

期刊

Complex and Intelligent Systems 封面图
Complex and Intelligent Systems
IF:
4.6
论文数:
2.1K
被引数:
6.6K

机构

A
Army Engineering University of PLA
学者数:
5.0K
论文数: 3.7K
被引数: 5
引用论文

引用论文

Imperceptible black-box waveform-level adversarial attack towards automatic speaker recognition
err2022-06-17
err7
errOAAI
errZhang, Xingyu; Zhang, Xiongwei; Sun, Meng; Zou, Xia; Chen, Kejiang; Yu, Nenghai
err分享
err收藏
Using Clinical Trial Simulators to Analyse the Sources of Variance in Clinical Trials of Novel Therapies for Acute Viral Infections
err2016-06-22
err0
errOAAI
errCarolin Vegvari; Emilie Cauët; Christoforos Hadjichrysanthou; Emma Lawrence; Gerrit-Jan Weverling; Frank de Wolf; Roy M. Anderson
err分享
err收藏
The effect of hydrogen on the magnetic properties of the Nd–Fe–B–H compound
err1999-05-01
err0
PREAI
errJ Tejada; J.M Hernandez; M Duran; E Krotenko; J.L Morenza; G Sardin; E.M Chudnovsky
err分享
err收藏
err分享
err收藏
One-shot voice conversion using a combination of U2-Net and vector quantization
err2022-10-01
err4
PREAI
errLiu, Fangkun; Wang, Hui; Ke, Yuxuan; Zheng, Chengshi
err分享
err收藏
学者 查看更多内容