arrow
Return

Voice conversion using General Regression Neural Network

delete2014-11-01
delete27
PRE
AI
J
Jagannath Nirmal *
M
Mukesh A. Zaveri
S
Suprava Patnaik
P
Pramod Kachare
DOI:10.1016/j.asoc.2014.06.040delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The objective of voice conversion system is to formulate the mapping function which can transform the source speaker characteristics to that of the target speaker. In this paper, we propose the General Regression Neural Network (GRNN) based model for voice conversion. It is a single pass learning network that makes the training procedure fast and comparatively less time consuming. The proposed system uses the shape of the vocal tract, the shape of the glottal pulse (excitation signal) and long term prosodic features to carry out the voice conversion task. In this paper, the shape of the vocal tract and the shape of source excitation of a particular speaker are represented using Line Spectral Frequencies (LSFs) and Linear Prediction (LP) residual respectively. GRNN is used to obtain the mapping function between the source and target speakers. The direct transformation of the time domain residual using Artificial Neural Network (ANN) causes phase change and generates artifacts in consecutive frames. In order to alleviate it, wavelet packet decomposed coefficients are used to characterize the excitation of the speech signal. The long term prosodic parameters namely, pitch contour (intonation) and the energy profile of the test signal are also modified in relation to that of the target (desired) speaker using the baseline method. The relative performances of the proposed model are compared to voice conversion system based on the state of the art RBF and GMM models using objective and subjective evaluation measures. The evaluation measures show that the proposed GRNN based voice conversion system performs slightly better than the state of the art models. (C) 2014 Elsevier B.V. All rights reserved.
Keywords:
Gaussian Mixture Model
General Regression Neural Network
Pitch contour
Radial Basis Function
Voice conversion
Wavelet transform
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

N
national institute of technology (nit system)
Scholars:
4.0W
Papers: 3.7W
Citations: 31
S
somaiya vidyavihar university
Scholars:
166
Papers: 175
Citations: 3
K
k j somaiya college of engineering
Scholars:
30
Papers: 28
Citations: 0
researcher View more organizations