arrow
Return

Self-generated Cross-Modal Prompt Tuning

delete2026-01-01
delete0
PRE
AI
G
Guiming Cao
W
Wu, Zonghan
H
Huan Huo
Y
Yuming Ou
G
Guandong Xu *
DOI:10.1007/978-3-032-06066-2_22delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Training prompt tuning models on task-specific data is a common method for adapting vision-language model knowledge to image recognition downstream tasks. Despite recent advancements in prompt tuning, achieving superior generalization to heterogeneous images, across a wide range of visual characteristics in style, format, and source, remains a significant challenge. To this end, we propose a novel method, namely Self-generated Cross-modal Prompt tuning (SCP), which generates pseudo prompts by applying the frozen knowledge in both the initialization and optimization stages to guide training. Consequently, the model can be trained on available datasets while effectively generalizing to heterogeneous image data in a wide spectrum of textual classes and visual characteristics. Extensive experiments on four benchmarks indicate that our proposed SCP significantly outperforms wellknown baselines in generalization performance across a broad spectrum of downstream tasks. Notably, our proposed SCP exhibits significant improvements in both Cross-Dataset and Domain-Shift Generalization, with performance gains of at least 3.63% and 11.71%, respectively. Our code is available at https://github.com/Ghosttimber/Academic.
Keywords:
Computer Vision
Multi-Modal
Prompt Tuning

Journal

M
MACHINE LEARNING AND KNOWLEDGE DISCOVERY IN DATABASES. RESEARCH TRACK, ECML PKDD 2025, PT III
IF:
0
Papers:
30
Citations:
0

Organization

E
east china normal university
Scholars:
3.1W
Papers: 2.1W
Citations: 25
U
university of technology sydney
Scholars:
1.6W
Papers: 2.0W
Citations: 25