Return
Cross-modal mapping: Mitigating the modality gap for few-shot classification
DOI:10.1016/j.patcog.2026.113285.png)
Abstract
En 中文
• A novel Cross-Modal Mapping method is proposed to mitigate the modality gap in vision-language models. • Linear transformation and triplet loss jointly optimize global and local cross-modal alignment. • CMM achieves a 1.06% average accuracy improvement across 11 datasets. • The method demonstrates excellent generalization on distribution shift datasets with high efficiency.
Keywords:
Cross-Modal Mapping
Modality Gap
Few-Shot Classification
Cross-Modal Alignment
Triplet Loss
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

