arrow
返回

Task-driven common subspace learning based semantic feature extraction for acoustic event recognition

delete2023-12-01
delete0
PRE
AI
邓世文 封面图
邓世文 (Shiwen Deng)
韩
韩纪庆 (Jiqing Han) *
DOI:10.1016/j.eswa.2023.121045delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
For acoustic event recognition (AER), it is important to extract the semantic feature that considers both the content information and the temporal ordering. To this end, our previous work proposed a common subspace learning (CSL) based method. However, the CSL treats the subspace learning and the back-end classifier training as two separate phases. In this manner, the discriminative information of the latter phase cannot be utilized to supervise the former learning. Therefore, the extracted feature based on the learned subspace also cannot contain the discriminative information. To solve this problem, we further propose a task-driven CSL (TD-CSL) based method to extract the semantic feature by jointly learning the above two phases. In the TD-CSL, the discriminative information, obtained during the classifier training phase, can be effectively adopted to supervise the learning of the common subspace. Specifically, the TD-CSL is formulated as a bi-level optimization problem, which regards the objective for the CSL in the lower level as a constraint of that for the classifier training in the upper. Furthermore, to obtain the optimal solutions of the subspace and the classifier, a gradient-based algorithm is designed. To evaluate the performance of the TD-CSL, experiments are conducted on the ESC-50 and ESC-10 databases. The TD-CSL can achieve 85.75% and 98.75% recognition accuracies on the two databases respectively, which outperforms the CSL and the related state-of-the-art methods.
Keyword:
Acoustic event recognition
Semantic feature
Discriminative information
Content information
Temporal ordering
Task-driven learning

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
H
Harbin Normal University
学者数:
4.1K
论文数: 2.5K
被引数: 3.8K
引用论文

引用论文

MTF-CRNN: Multiscale Time-Frequency Convolutional Recurrent Neural Network for Sound Event Detection
err2020-01-01
err14
errOAAI
errZhang, Keming; Cai, Yuanwen; Ren, Yuan; Ye, Ruida; He, Liang
err分享
err收藏
err分享
err收藏
Deep Belief Network based audio classification for construction sites monitoring
err2021-09-01
err30
errOAAI
errScarpiniti, Michele; Colasante, Francesco; Di Tanna, Simone; Ciancia, Marco; Lee, Yong-Cheol; Uncini, Aurelio
err分享
err收藏
Masked Conditional Neural Networks for sound classification
err2020-05-01
err36
errOAAI
errMedhat, Fady; Chesmore, David; Robinson, John
err分享
err收藏
学者 查看更多内容