arrow
返回

Generating visual-adaptive audio representation for audio recognition

delete2025-06-01
delete0
PRE
AI
J
J. Youn
D
Dae Ung Jo
S
Seung Mo Seo
S
S.H. Kim
J
Jongwon Choi *
DOI:10.1016/j.patrec.2025.03.020delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We propose Visual-adaptive Audio Spectrogram Generation (VASG), which is an innovative audio feature generation method preserving the Mel-spectrogram's structure while enhancing its own discriminability. VASG maintains the spatio-temporal information of the Mel-spectrogram without degrading the performance of existing audio recognition and improves intra-class discriminability by incorporating the relational knowledge of images. VASG incorporates images only during the training phase, and once trained, VASG can be utilized as a converter that takes an input Mel-spectrogram and outputs an enhanced Mel-spectrogram, improving the discriminability of audio spectrograms without requiring further training during application. To effectively increase the discriminability of the encoded audio feature, we introduce a novel audio-visual correlation learning loss, named Batch-wise Correlation Transfer loss, that aligns inter-correlation between audio and visual modality. When applying pre-trained VASG to convert environmental sound classification benchmarks, we observed performance improvements in various audio classification models. Using the enhanced Mel-spectrograms produced by VASG, as opposed to the original Mel-spectrogram input, led to performance gains in recent state-of-the-art models, with accuracy increases of up to 4.27%.
Keyword:
Multimodal learning
Audiovisual learning
Contrastive learning
Audio classification

期刊

Pattern Recognition Letters 封面图
Pattern Recognition Letters
IF:
3.3
论文数:
8.0K
被引数:
1.6W

机构

C
Chung Ang University
学者数:
1.3W
论文数: 1.4W
被引数: 133
K
kyungpook national university (knu)
学者数:
1.8W
论文数: 1.8W
被引数: 14
引用论文

引用论文

暂无论文信息