arrow
返回

Audio classification using attention-augmented convolutional neural network

delete2018-12-01
delete32
delete
OA
AI
Y
Yu Wu
H
Hua Mao *
Y
Yi Zhang
DOI:10.1016/j.knosys.2018.07.033delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Audio classification, as a set of important and challenging tasks, groups speech signals according to speakers' identities, accents, and emotional states. Due to the high dimensionality of the audio data, task-specific hand-crafted features extraction is always required and regarded cumbersome for various audio classification tasks. More importantly, the inherent relationship among features has not been fully exploited. In this paper, the original speech signal is first represented as spectrogram and later be split along the frequency domain to form frequency-distributed spectrogram. This paper proposes a task-independent model, called FreqCNN, to automaticly extract distinctive features from each frequency band by using convolutional kernels. Further more, an attention mechanism is introduced to systematically enhance the features from certain frequency bands. The proposed FreqCNN is evaluated on three publicly available speech databases thorough three independent classification tasks. The obtained results demonstrate superior performance over the state-of-the-art.
Keyword:
Audio classification
Spectrograms
Convolutional neural networks
Attention mechanism
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

S
sichuan university
学者数:
12.1W
论文数: 7.8W
被引数: 100
引用论文

引用论文

err分享
err收藏
Speech emotion recognition using amplitude modulation parameters and a combined feature selection procedure
err2014-06-01
err68
PREAI
errMencattini, Arianna; Martinelli, Eugenio; Costantini, Giovanni; Todisco, Massimiliano; Basile, Barbara; Bozzali, Marco; Di Natale, Corrado
err分享
err收藏
Limiting oxygen partial pressure of DyMnO3 phase
err1984-05-01
err0
PREAI
errNaoki Kamegashira; Yoshihiko Hiyoshi
err分享
err收藏
Toxic megacolon in ulcerative colitis complicated by pneumomediastinum: Report of two cases
err1980-09-01
err0
errOAAI
errGlen R. Mogan; David B. Sachar; Joel Bauer; Barry Salky; Henry D. Janowitz
err分享
err收藏
Employing unlabeled data to improve the classification performance of SVM, and its application in audio event classification
err2016-04-01
err15
PREAI
errLeng, Yan; Sun, Chengli; Xu, Xinyan; Yuan, Qi; Xing, Shuning; Wan, Honglin; Wang, Jingjing; Li, Dengwang
err分享
err收藏
Deep neural network framework and transformed MFCCs for speaker's age and gender classification
err2017-01-01
err54
PREAI
errQawaqneh, Zakariya; Abu Mallouh, Arafat; Barkana, Buket D.
err分享
err收藏
学者 查看更多内容