arrow
Return

Optimizing Multi-Taper Features for Deep Speaker Verification

delete2021-01-01
delete0
delete
OA
AI
X
Xuechen Liu *
S
Sahidullah, Md
T
Tomi Kinnunen
DOI:10.1109/LSP.2021.3122796delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multi-taper estimators provide low-variance power spectrum estimates that can be used in place of the windowed discrete Fourier transform (DFT) to extract speech features such as mel-frequency cepstral coefficients (MFCCs). Even if past work has reported promising automatic speaker verification (ASV) results with Gaussian mixture model-based classifiers, the performance of multi-taper MFCCs with deep ASV systems remains an open question. Instead of a static-taper design, we propose to optimize the multi-taper estimator jointly with a deep neural network trained for ASV tasks. With a maximum improvement on the SITW corpus of 25.8% in terms of equal error rate over the static-taper, our method helps preserve a balanced level of leakage and variance, providing more robustness.
Keywords:
Feature extraction
Discrete Fourier transforms
Task analysis
Neural networks
Mel frequency cepstral coefficient
Stochastic processes
Standards
Multi-taper spectrum
speaker verification

Journal

IEEE Signal Processing Magazine cover
IEEE Signal Processing Magazine
IF:
9.6
Papers:
1.1W
Citations:
1.7W

Organization

C
centre national de la recherche scientifique (cnrs)
Scholars:
24.5W
Papers: 18.2W
Citations: 279
I
Inria
Scholars:
3.5K
Papers: 2.5K
Citations: 343