arrow
返回

Mel-Weighted Single Frequency Filtering Spectrogram for Dialect Identification

delete2020-01-01
delete11
delete
OA
AI
R
Rashmi Kethireddy *
S
Sudarsana Reddy Kadiri
P
Paavo Alku
S
Suryakanth V. Gangashetty
DOI:10.1109/ACCESS.2020.3020506delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
In this study, we propose Mel-weighted single frequency filtering (SFF) spectrograms for dialect identification. The spectrum derived using SFF has high spectral resolution for harmonics and resonances while simultaneously maintaining good time-resolution of some speech excitation features such as impulse-like events. The SFF spectrum can represent speech characteristics such as burst time and glottal closure instances better than the short-time Fourier transform (STFT) spectrum. Our hypothesis is that these intricate representations in the SFF spectrum should help in distinguishing dialects. Therefore, we built a dialect identification system which uses an unsupervised, bottleneck feature representation of the Mel-weighted SFF spectrogram (Mel-SFF spectrogram) with sequence-to-sequence deep autoencoders. The language invariance of the proposed system was evaluated using two datasets: the UT-Podcast database (English) and the STYRIALECT database (German). The proposed representations gave a relative improvement of 9.47% and 4.69% in unweighted average recall (UAR) compared to the best baseline method on the development and test datasets, respectively, of the UT-Podcast database. The proposed representations also gave a comparable performance to the best baseline method for the STYRIALECT database. In addition, the fusion of the autoencoder bottleneck features computed from the Mel-SFF and Mel-STFT spectrograms improved the overall performance indicating complementary information between these features. By further analyzing the performance of the proposed representation with different utterance lengths using the UT-Podcast database, we observed that the proposed representation performed better on short utterances. The improved performance given by the Mel-weighted SFF spectrogram for recognizing dialects in both databases supports our hypothesis.
Keyword:
Spectrogram
Databases
Acoustics
Resonant frequency
Speech processing
Harmonic analysis
Licenses
Dialect identification
single frequency filtering (SFF) spectrum
Mel-spectrogram
Mel-filter bank energies
autoencoder
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

A
Aalto University
学者数:
1.6W
论文数: 1.5W
被引数: 2.1W
I
International Institute of Information Technology Hyderabad
学者数:
779
论文数: 685
被引数: 5
引用论文

引用论文

INTRINSIC XASE LIGAND INTERACTIONS IMPACT FVIIIA REGULATION内在Xase配体相互作用影响FVIIIA的调控
err2024-12-01
err0
PREAI
errMorris, John J.; Parsons, Nicole A.; Davidson, Robert J.; George, Lindsey A.
err分享
err收藏
err分享
err收藏
i-Vector Modeling of Speech Attributes for Automatic Foreign Accent Recognition
err2016-01-01
err30
PREAI
errBehravan, Hamid; Hautamaki, Ville; Siniscalchi, Sabato Marco; Kinnunen, Tomi; Lee, Chin-Hui
err分享
err收藏
学者 查看更多内容