Return
Audio-Visual Event Localization With Cross Co-Attention and Dynamic Audio-Object Semantic Alignment
DOI:10.1109/LSP.2025.3589935.png)
Abstract
En 中文
In Audio-Visual Event Localization (AVEL) task, various cross-modal attentions (CMA) were proposed to capture the bilateral correlations of audio and visual segments. However, existing CMA approaches are inefficient since they require two sets of independent attention parameters. Besides, existing works often ignore the semantic alignment between audio and audible objects, leading to the suboptimal localization results. In this letter, a novel network with a cross co-attention (CCA) and a dynamic audio-object semantic alignment (DAOSA) strategy is proposed to tackle these issues. Unlike existing CMA methods, CCA calculates the co-attention between audio and visual segments to capture the bilateral correlations via a group of parameters. To align the semantics of audio and audio-related objects, DAOSA proposes a dynamic threshold scheme to adaptively select the highly relevant audio-object pairs as positivity while regarding other pairs as negativity. Then, DAOSA optimizes the semantic alignment of positive pairs by contrastive learning. Experiments across different datasets demonstrate the effectiveness of proposed method, which also outperforms several state-of-the-art models.
Keywords:
Audio-visual event localization (AVEL)
attention mechanism
semantic alignment (SA)
Journal
IF:
9.6
Papers:
1.1W
Citations:
1.7W

