Return
Multi-feature wavelet attention network for audio deepfake detection
王
R
王
Z
DOI:10.1016/j.knosys.2026.116757.png)
Abstract
En 中文
• Acoustic and emotional features are fused for more comprehensive speech representation. • Wavelet transform is integrated into deep neural networks to extract detailed speech features. • A multi-scale attention mechanism is used to enhance detection performance. • The method achieves 0.18% EER on ASVspoof 2019 LA and 2.55% on ASVspoof 2021 DF dataset. • Superior generalization is shown with 9.08% EER on the In-the-Wild dataset.
Keywords:
Wavelet transform
Automatic speaker verification
Audio deepfake detection
Self-supervised learning
Wav2Vec 2.0 XLSR
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
