arrow
Return

Multimodal Voice Conversion Under Adverse Environment Using a Deep Convolutional Neural Network

delete2019-01-01
delete1
delete
OA
AI
周健 (Jian Zhou) *
Y
Yuting Hu
H
Hailun Lian
H
Huabin Wang
陶亮 (Liang Tao)
H
Hon Keung Kwan
DOI:10.1109/ACCESS.2019.2955982delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This paper presents a voice conversion (VC) technique under noisy environments. Typically, VC methods use only audio information for conversion in a noiseless environment. However, existing conversion methods do not always achieve satisfactory results in an adverse acoustic environment. To solve this problem, we propose a multimodal voice conversion model based on a deep convolutional neural network (MDCNN) built by combining two convolutional neural networks (CNN) and a deep neural network (DNN) for VC under noisy environments. In the MDCNN, both the acoustic and visual information are incorporated into the voice conversion to improve its robustness in adverse acoustic conditions. The two CNNs are designed to extract acoustic and visual features, and the DNN is designed to capture the nonlinear mapping relation of source speech and target speech. Experimental results indicate that the proposed MDCNN outperforms two existing approaches in noisy environments.
Keywords:
Feature extraction
Visualization
Acoustics
Noise measurement
Lips
Convolutional neural nets
Audio and video feature fusion
convolutional neural network
deep learning
mel-frequency cepstral coefficients
multilayer feedforward neural networks
multimodal voice conversion
noise robustness
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
university of windsor
Scholars:
4.4K
Papers: 4.5K
Citations: 3
A
anhui university
Scholars:
1.9W
Papers: 1.2W
Citations: 24