arrow
Return

Nonparallel Voice Conversion With Augmented Classifier Star Generative Adversarial Networks

delete2020-01-01
delete15
delete
OA
AI
H
Hirokazu Kameoka *
T
Takuhiro Kaneko
K
Kou Tanaka
N
Nobukatsu Hojo
DOI:10.1109/TASLP.2020.3036784delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We previously proposed a method that allows for nonparallel voice conversion (VC) by using a variant of generative adversarial networks (GANs) called StarGAN. The main features of our method, called StarGAN-VC, are as follows: First, it requires no parallel utterances, transcriptions, or time alignment procedures for speech generator training. Second, it can simultaneously learn mappings across multiple domains using a single generator network and thus fully exploit available training data collected from multiple domains to capture latent features that are common to all the domains. Third, it can generate converted speech signals quickly enough to allow real-time implementations and requires only several minutes of training examples to generate reasonably realistic-sounding speech. In this article, we describe three formulations of StarGAN, including a newly introduced novel StarGAN variant called Augmented classifier StarGAN (A-StarGAN), and compare them in a nonparallel VC task. We also compare them with several baseline methods.
Keywords:
Voice conversion (VC)
nonparallel VC
multi-domain VC
generative adversarial networks (GANs)
CycleGAN
StarGAN
A-StarGAN
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

No organization information available