Return
Multimodal Domain Generalization via Adaptive Dual-Objective Feature Learning
H
S
G
S
M
DOI:10.1007/s11263-026-02991-0.png)
Abstract
En 中文
Recent years have witnessed the widespread adoption of multimodal applications across various real-world scenarios. However, these applications often face significant challenges due to distribution shifts across different domains, which can severely impact model performance and reliability. While domain generalization has been extensively studied for unimodal scenarios, its multimodal counterpart remains underexplored, presenting unique challenges in addressing both modality heterogeneity and domain shifts simultaneously. We argue that an effective multimodal domain generalization framework should not only learn domain invariance but also balance the invariance and specificity of modalities. To bridge this gap, we introduce Adaptive Dual-Objective Feature Learning for Multimodal Domain Generalization (ADMMDG), a framework that separates feature learning into two complementary components: features that capture shared patterns with both modality and domain invariance, and modality specific features that preserve the unique characteristics of each modality while maintaining domain invariance. To achieve this, ADMMDG employs two distinct contrastive learning strategies that focus on aligning features across different domains and modalities. Both strategies are equipped with a dynamic weighting mechanism that assigns higher weights to modal-domain feature pairs exhibiting greater modality or domain discrepancies, thereby effectively capturing the respective feature learning objectives. Extensive experiments on multiple benchmark datasets, utilizing diverse modality combinations, demonstrate the effectiveness of ADMMDG in both multi-source and single-source multimodal domain generalization tasks, as well as in the challenging multimodal open-set domain generalization scenario. Our codes are available at https://github.com/lihongzhao99/ADMMDG .
Keywords:
Domain Generalization
Multimodal
Adaptive Feature Learning
Contrastive Learning
Journal
IF:
9.3
Papers:
3.9K
Citations:
2.8W
