arrow
返回

Attention-Based Temporal-Frequency Aggregation for Speaker Verification

delete2022-03-10
delete5
delete
OA
AI
M
Meng Wang
D
Da‐Zheng Feng *
T
Tingting Su
陈默涵 封面图
陈默涵 (Mohan Chen)
DOI:10.3390/s22062147delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Convolutional neural networks (CNNs) have significantly promoted the development of speaker verification (SV) systems because of their powerful deep feature learning capability. In CNN-based SV systems, utterance-level aggregation is an important component, and it compresses the frame-level features generated by the CNN frontend into an utterance-level representation. However, most of the existing aggregation methods aggregate the extracted features across time and cannot capture the speaker-dependent information contained in the frequency domain. To handle this problem, this paper proposes a novel attention-based frequency aggregation method, which focuses on the key frequency bands that provide more information for utterance-level representation. Meanwhile, two more effective temporal-frequency aggregation methods are proposed in combination with the existing temporal aggregation methods. The two proposed methods can capture the speaker-dependent information contained in both the time domain and frequency domain of frame-level features, thus improving the discriminability of speaker embedding. Besides, a powerful CNN-based SV system is developed and evaluated on the TIMIT and Voxceleb datasets. The experimental results indicate that the CNN-based SV system using the temporal-frequency aggregation method achieves a superior equal error rate of 5.96% on Voxceleb compared with the state-of-the-art baseline models.
Keyword:
convolutional neural networks
speaker verification
temporal-frequency aggregation
self-attention
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Sensors 封面图
Sensors
IF:
3.5
论文数:
7.2W
被引数:
20.9W

机构

X
Xidian University
学者数:
2.4W
论文数: 1.9W
被引数: 9.7K
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Dilated residual networks with multi-level attention for speaker verification
err2020-10-01
err16
PREAI
errWu, Yanfeng; Guo, Chenkai; Gao, Hongcan; Xu, Jing; Bai, Guangdong
err分享
err收藏
Speaker Recognition by Machines and Humans
err2015-11-01
err464
PREAI
errHansen, John H. L.; Hasan, Taufiq
err分享
err收藏
Self-attention based speaker recognition using Cluster-Range Loss
err2019-11-01
err19
PREAI
errBian, Tengyue; Chen, Fangzhou; Xu, Li
err分享
err收藏
Planning future studies based on the conditional power of a meta‐analysis
err2012-07-11
err0
errOAAI
errVerena Roloff; Julian P.T. Higgins; Alex J. Sutton
err分享
err收藏
Tumorigenicity of Methyl-n-Propylnitrosamine in Syrian Golden Hamsters2
err1974-02-01
err0
PREAI
errParviz Pour; Friedrich W. Krüger; Antonio Cardesa; Jürgen Althoff; Ulrich Mohr
err分享
err收藏
学者 查看更多内容