返回
Data Augmentation Using Deep Generative Models for Embedding Based Speaker Recognition
DOI:10.1109/TASLP.2020.3016498.png)
摘要
En 中文
Data augmentation is an effective method to improve the robustness of embedding based speaker verification systems, which could be applied to either the front-end speaker embedding extractor or the back-end PLDA. Different from the conventional augmentation methods such as manually adding noise or reverberation to the original audios, in this article, we propose to use deep generative models to directly generate more diverse speaker embeddings, which would be used for robust PLDA training. Conditional GAN, and VAE are designed, and investigated for different embedding types, including factor analysis based i-vector, TDNN based x-vector, and ResNet based r-vector. The proposed back-end augmentation methods are evaluated on NIST SRE 2016, and 2018 dataset. Within the popular x-vector, and r-vector framework, the experimental results show that our proposed methods can outperform the traditional audio based back-end augmentation method while different front-end augmentation methods are considered.
Keyword:
Generative adversarial networks
Data mining
Speaker recognition
Gallium nitride
Feature extraction
Training
Speech recognition
Text-independent speaker verification
data augmentation
generative adversarial network
variational auto-encoder
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
5.1
论文数:
2.6K
被引数:
1.1W
机构
引用论文
Pequi (Caryocar coriaceum Wittm., Caryocaraceae) Oil Production: A strong economically influenced tradition in the Araripe region, northeastern BrazilPequi (Caryocar coriaceum Wittm., Caryocaraceae)油的生产:一种在经济上受到强烈影响的传统,位于巴西东北部的阿腊里皮地区。
Investigation of sprout-growth-inhibitory compounds in the volatile fraction of potato tubers马铃薯块茎挥发性组分中发芽抑制化合物的调查研究
Validating lactate dehydrogenase (LDH) as a component of the PLASMIC predictive tool (PLASMIC-LDH)验证乳酸脱氢酶(LDH)作为PLASMIC预测工具(PLASMIC-LDH)的组成部分

