Return
AVQACL++: Toward a Robust Framework and Benchmark for Audio-Visual Question Answering Continual Learning
DOI:10.1109/TCSVT.2026.3662386.png)
Abstract
En 中文
This paper presents a unified framework for continual learning in Audio-Visual Question Answering (AVQACL), designed to enhance fine-grained scene understanding and spatio-temporal reasoning in dynamic multimodal task environments. To simulate realistic incremental learning scenarios, we construct two large-scale benchmark datasets, Split-AVQA and Split-MUSIC-AVQA, by reorganizing existing AVQA corpora into sequential tasks. Empirical results show that conventional models suffer from severe performance degradation and catastrophic forgetting when learning across audio, visual, and textual modalities in a continual setup. To address these challenges, we propose a novel approach that integrates four key modules: 1) Question-Guided Cross-modal Information Fusion (QCIF), which dynamically extracts task-relevant multimodal features via question-aware attention; 2) Task-specific Knowledge Distillation with Spatial-Temporal Feature Constraints (TKD-STFC), which preserves semantic output behavior and internal reasoning trajectories across tasks; 3) Question Semantic Consistency Constraint (QSCC), which regularizes evolving question representations to maintain linguistic stability; and 4) Dual-Strategy Exemplar Selection (DSES), a memory-efficient replay strategy that jointly maximizes sample informativeness and diversity. All components are theoretically grounded, and formal analysis is provided to ensure modality alignment, spatial-temporal coherence, and exemplar selection reliability. Extensive experiments on both datasets demonstrate that our method consistently outperforms prior state-of-the-art AVQACL baselines in terms of accuracy, retention, and robustness.
Keywords:
Multimodal continual learning
audio-visual question answering
video scene understanding
spatial-temporal reasoning
catastrophic forgetting
Journal
IF:
11.1
Papers:
612
Citations:
3.1W

