返回
Task-driven common subspace learning based semantic feature extraction for acoustic event recognition
DOI:10.1016/j.eswa.2023.121045.png)
摘要
En 中文
For acoustic event recognition (AER), it is important to extract the semantic feature that considers both the content information and the temporal ordering. To this end, our previous work proposed a common subspace learning (CSL) based method. However, the CSL treats the subspace learning and the back-end classifier training as two separate phases. In this manner, the discriminative information of the latter phase cannot be utilized to supervise the former learning. Therefore, the extracted feature based on the learned subspace also cannot contain the discriminative information. To solve this problem, we further propose a task-driven CSL (TD-CSL) based method to extract the semantic feature by jointly learning the above two phases. In the TD-CSL, the discriminative information, obtained during the classifier training phase, can be effectively adopted to supervise the learning of the common subspace. Specifically, the TD-CSL is formulated as a bi-level optimization problem, which regards the objective for the CSL in the lower level as a constraint of that for the classifier training in the upper. Furthermore, to obtain the optimal solutions of the subspace and the classifier, a gradient-based algorithm is designed. To evaluate the performance of the TD-CSL, experiments are conducted on the ESC-50 and ESC-10 databases. The TD-CSL can achieve 85.75% and 98.75% recognition accuracies on the two databases respectively, which outperforms the CSL and the related state-of-the-art methods.
Keyword:
Acoustic event recognition
Semantic feature
Discriminative information
Content information
Temporal ordering
Task-driven learning
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
MTF-CRNN: Multiscale Time-Frequency Convolutional Recurrent Neural Network for Sound Event Detection
IEEE ACCESS
IF3.6
A new pyramidal concatenated CNN approach for environmental sound classification一种新的金字塔级联CNN环境声音分类方法
APPLIED ACOUSTICS
IF3.6
Multimodal fusion methods with deep neural networks and meta-information for aggression detection in surveillance具有深度神经网络和元信息的多模式融合方法,用于监视中的攻击检测
Adv-ESC: Adversarial attack datasets for an environmental sound classificationAdv-esc: 环境声音分类的对抗攻击数据集
APPLIED ACOUSTICS
IF3.6

