arrow
Return

Dynamic interactive learning network for audio-visual event localization

delete2023-11-18
delete0
PRE
AI
J
Jincai Chen
H
Han Liang
R
Ruili Wang
J
Jiangfeng Zeng *
卢萍 cover
卢萍 (Ping Lü)
DOI:10.1007/s10489-023-05146-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Audio-visual event (AVE) localization aims to detect whether an event exists in each video segment and predict its category. Only when the event is audible and visible can it be recognized as an AVE. However, sometimes the information from auditory and visual modalities is asymmetrical in a video sequence, leading to incorrect predictions. To address this challenge, we introduce a dynamic interactive learning network designed to dynamically explore the intra- and inter-modal relationships depending on the other modality for better AVE localization. Specifically, our approach involves a dynamic fusion attention of intra- and inter-modalities module, enabling the auditory and visual modalities to focus more on regions deemed informative by the other modality while focusing less on regions that the other modality considers noise. In addition, we introduce an audio-visual difference loss to reduce the distance between auditory and visual representations. Our proposed method has been demonstrated to have superior performance by extensive experimental results on the AVE dataset. The source code will be available at https://github.com/hanliang/DILN.
Keywords:
Audio-visual event localization
Dynamic fusion
Attention mechanism
Difference loss

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.5K
Citations:
1.7W

Organization

C
Central China Normal University
Scholars:
1.1W
Papers: 8.1K
Citations: 1.1W
M
Massey University
Scholars:
7.6K
Papers: 7.8K
Citations: 9.6K