arrow
Return

Statistical Parametric Speech Synthesis Based on Gaussian Process Regression

delete2014-04-01
delete27
delete
OA
AI
T
Tomoki Koriyama *
T
Takashi Nose
T
Takao Kobayashi
DOI:10.1109/JSTSP.2013.2283461delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper proposes a statistical parametric speech synthesis technique based on Gaussian process regression (GPR). The GPR model is designed for directly predicting frame-level acoustic features from corresponding information on frame context that is obtained from linguistic information. The frame context includes the relative position of the current frame within the phone and articulatory information and is used as the explanatory variable in GPR. Here, we introduce cluster-based sparse Gaussian processes (GPs), i.e., local GPs and partially independent conditional (PIC) approximation, to reduce the computational cost. The experimental results for both isolated phone synthesis and full-sentence continuous speech synthesis revealed that the proposed GPR-based technique without dynamic features slightly outperformed the conventional hidden Markov model (HMM)-based speech synthesis using minimum generation error training with dynamic features.
Keywords:
Gaussian process regression
nonparametric Bayesian model
partially independent conditional (PIC) approximation
sparse Gaussian processes
statistical speech synthesis

Journal

IEEE Journal of Selected Topics in Signal Processing cover
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
Papers:
1.9K
Citations:
1.1W

Organization

I
Institute of Science Tokyo
Scholars:
3.2W
Papers: 2.7W
Citations: 117