返回
Speech activity detection using time-frequency auditory spectral pattern
DOI:10.1016/j.apacoust.2020.107403.png)
摘要
En 中文
Speech activity detection (SAD) is useful to identify human speech from other sounds and ear plays an important role in this. Gammatone filter bank response mimics the human auditory periphery model. We propose an unsupervised SAD based on cochleagram features which are derived from gammatone filter bank-multiple window observation (GTFB-MWO) technique. Spectral clustering of cochleagram features is employed here with two strategies: correlation measure based non-speech input detection and feature winning score to identify the speech and non-speech clusters. The proposed SAD performance is compared with different baseline unsupervised SAD techniques on 600 min of noisy speech corpus over different signal-to-noise ratios (SNRs). The result shows 4.52% increase in overall F1 - score and 3.42% reduction in half total error rate (HTER) compared to the recent reported result of rVAD [1]. (C) 2020 Elsevier Ltd. All rights reserved.
Keyword:
Cochleagram
Gammatone filter
Spectral clustering
Speech activity detection
Unsupervised learning
期刊
IF:
3.6
论文数:
7.3K
被引数:
1.7W

