arrow
Return

How does Layer Normalization improve Batch Normalization in self-supervised sound source localization?

delete2024-01-01
delete2
PRE
AI
T
Tianyu Liu
张鹏 cover
张鹏 (Peng Zhang) *
黄维 cover
黄维 (Wei Huang)
Y
Yufei Zha
T
Tao You
Y
Yanning Zhang
DOI:10.1016/j.neucom.2023.127040delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Self-supervised sound source localization is usually challenged by the unexpected large input and incorrect direction of normalization in current solutions. A promising way for this challenge is to avoid feature deformation by incorporating more effective normalization, which is the motivation of this study. Based on the mathematical derivation of Layer Normalization (LN) in scale independence, in this work, a correspondence consolidation method is proposed to reinforce the audio-visual correspondence. By ensembling input feature normalization and LN-based simsiam Predictor, a joint gradient stabilization can be further achieved for more accurate sound source localization. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have verified a superior performance in comparison to the other state-of-the-art works.
Keywords:
Audio-visual
Sound source localization
Batch Normalization
Layer Normalization

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

N
Nanchang University
Scholars:
3.7W
Papers: 2.1W
Citations: 3.7W
N
Northwestern Polytechnical University
Scholars:
4.6W
Papers: 3.7W
Citations: 5.3W