arrow
Return

Cross-modal mapping: Mitigating the modality gap for few-shot classification

delete2026-02-10
delete0
PRE
AI
X
Xi Yang
W
Wulin Xie
P
Pai Peng
文杰 (Jie Wen)
卢晓寰 cover
卢晓寰 (Xiaohuan Lu)
DOI:10.1016/j.patcog.2026.113285delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• A novel Cross-Modal Mapping method is proposed to mitigate the modality gap in vision-language models. • Linear transformation and triplet loss jointly optimize global and local cross-modal alignment. • CMM achieves a 1.06% average accuracy improvement across 11 datasets. • The method demonstrates excellent generalization on distribution shift datasets with high efficiency.
Keywords:
Cross-Modal Mapping
Modality Gap
Few-Shot Classification
Cross-Modal Alignment
Triplet Loss

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

H
Harbin Institute of Technology
Scholars:
1.5W
Papers: 4.7K
Citations: 8.5W
G
Guizhou University
Scholars:
3.3K
Papers: 1.1K
Citations: 1.6W