arrow
返回

Robust Speaker Verification Using Deep Weight Space Ensemble

delete2023-01-01
delete2
PRE
AI
林伟伟 封面图
林伟伟 (Weiwei Lin)
M
Man‐Wai Mak *
DOI:10.1109/TASLP.2022.3233231delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Domain shift is one of the most challenging problems in speaker verification. Although numerous methods have been proposed to address domain shift, most approaches optimize the performance of one domain at the sacrifice of the other. As a result, to obtain the best performance, each domain requires a dedicated model. However, deploying multiple models is resource-demanding and impractical, particularly when the deployment domains are not known in advance. Recent studies in deep neural networks (DNNs) suggest that near the low error surface of the DNN's weight space, there exists a linear path connecting a base model and a fine-tuned model. This finding inspires us to combine the strength of the fine-tuned models and the base models to solve challenging SV problems. Specifically, we aim to develop models that can handle 1) mixed text-dependent (TD) and text-independent (TI) speaker verification where the speech content can be either unconstrained or constrained, 2) cross-channel speaker verification where the recording can be 16 kHz high-fidelity microphone speech or 8 kHz telephone speech, and 3) bi-lingual speaker verification where the enrollment and test speech can be one of the two languages. With weight space ensemble, we show that we can substantially improve the tasks mentioned above, with a 39.6% improvement in mixing TD and TI SV, a 17.4% improvement in bi-lingual SV, and an 18.4% improvement in cross-channel SV. Moreover, we show that the weight space ensemble can also enhance the performance in the target domain, thanks to the regularization effect of the interpolation.
Keyword:
Training
Data models
Kernel
Convolution
Computer architecture
Adaptation models
Telephone sets
Robust speaker recognition
domain adaptation
domain shift
weight space ensemble

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

H
hong kong polytechnic university
学者数:
3.0W
论文数: 4.1W
被引数: 921
引用论文

引用论文

Planning future studies based on the conditional power of a meta‐analysis
err2012-07-11
err0
errOAAI
errVerena Roloff; Julian P.T. Higgins; Alex J. Sutton
err分享
err收藏
Tumorigenicity of Methyl-n-Propylnitrosamine in Syrian Golden Hamsters2
err1974-02-01
err0
PREAI
errParviz Pour; Friedrich W. Krüger; Antonio Cardesa; Jürgen Althoff; Ulrich Mohr
err分享
err收藏
Mechanical performance of thermoplastic olefin composites reinforced with coir and sisal natural fibers: Influence of surface pretreatment
err2019-01-25
err0
PREAI
errJoão F. Pereira; Diana P. Ferreira; João Bessa; Joana Matos; Fernando Cunha; Isabel Araújo; Luís F. Silva; Elizabete Pinho; Raul Fangueiro
err分享
err收藏
Seasonal and elevational variation in glucose and glycogen in two songbird species
err2020-07-01
err0
PREAI
errKaren L. Sweazea; Krystal S. Tsosie; Elizabeth J. Beckman; Phred M. Benham; Christopher C. Witt
err分享
err收藏
Southeast Asian diversity: first insights into the complex mtDNA structure of Laos
err2011-02-18
err0
errOAAI
errMartin Bodner; Bettina Zimmermann; Alexander Röck; Anita Kloss-Brandstätter; David Horst; Basil Horst; Sourideth Sengchanh; Torpong Sanguansermsri; Jürgen Horst; Tanja Krämer; Peter M Schneider; Walther Parson
err分享
err收藏
学者 查看更多内容