返回
Generating visual-adaptive audio representation for audio recognition
DOI:10.1016/j.patrec.2025.03.020.png)
摘要
En 中文
We propose Visual-adaptive Audio Spectrogram Generation (VASG), which is an innovative audio feature generation method preserving the Mel-spectrogram's structure while enhancing its own discriminability. VASG maintains the spatio-temporal information of the Mel-spectrogram without degrading the performance of existing audio recognition and improves intra-class discriminability by incorporating the relational knowledge of images. VASG incorporates images only during the training phase, and once trained, VASG can be utilized as a converter that takes an input Mel-spectrogram and outputs an enhanced Mel-spectrogram, improving the discriminability of audio spectrograms without requiring further training during application. To effectively increase the discriminability of the encoded audio feature, we introduce a novel audio-visual correlation learning loss, named Batch-wise Correlation Transfer loss, that aligns inter-correlation between audio and visual modality. When applying pre-trained VASG to convert environmental sound classification benchmarks, we observed performance improvements in various audio classification models. Using the enhanced Mel-spectrograms produced by VASG, as opposed to the original Mel-spectrogram input, led to performance gains in recent state-of-the-art models, with accuracy increases of up to 4.27%.
Keyword:
Multimodal learning
Audiovisual learning
Contrastive learning
Audio classification
期刊
IF:
3.3
论文数:
8.0K
被引数:
1.6W
机构
引用论文
暂无论文信息

