Return
CAMDG: A complementary multilevel alignment framework for multimodal domain generalization
DOI:10.1016/j.neucom.2026.134560.png)
Abstract
En 中文
Domain generalization (DG) aims to learn transferable knowledge from known source domains to maintain robust generalization performance on unseen target domains with different distributions. Although progress has been made in single-modal DG tasks, multimodal domain generalization (MMDG) still faces major challenges. The interaction between domain distribution discrepancies and modality distribution discrepancies makes it difficult to learn consistent cross modal and cross domain feature representations, thus limiting generalization performance. In this paper, we propose a complementary alignment framework for multimodal domain generalization (CAMDG). Specifically, we employ a multisource domain alignment (MSDA) module to align multiple source domains in a shared feature space. We further introduce a cross modal feature alignment (CMFA) module to align distributional discrepancies among modalities, and a cross modal rationale alignment (CMRA) module to constrain the decision-level rationale consistency across modalities. Extensive experiments on the EPIC-Kitchens and Human–Animal–Cartoon (HAC) datasets demonstrate that our proposed framework consistently improves cross modal domain generalization capability with the average accuracy exceeding state-of-the-art methods by 1.9% and 2.02%, respectively.

