arrow
Return

Multi-modal recursive prompt learning with mixup embedding for generalization recognition

delete2024-06-01
delete2
PRE
AI
Y
Yunpeng Jia
X
Xiufen Ye *
刘育松 (Yusong Liu)
郭书祥 (Shuxiang Guo)
DOI:10.1016/j.knosys.2024.111726delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The contrastive language-image pretraining (CLIP) model has shown promise in generalization recognition by combining visual and textual embeddings. However, the pretrained CLIP model requires further fine tuning for downstream tasks. Existing prompt learning (PL) approaches, while effective, neglect the cross-hierarchical fusion of multimodal features. To address this issue, we introduce multi -modal recursive PL (MmRPL) for different generalization image recognition tasks. Recursive connections between prompts at different layers facilitate hierarchy-aware vision -text fusion. To our best knowledge, we introduce a mixup embedding technique to enhance feature representations of PL for the first time. Extensive experiments across popular generalization recognition settings reveal that MmRPL outperforms existing methods. Furthermore, when combined with the embedding technique, MmRPL further enhances recognition in base-to-novel generalization, generalized zeroshot, domain generalization, and domain adaptive zero-shot scenarios. Notably, the proposed methods extend to applications like zero-shot recognition of side-scan sonar (SSS) image targets via domain generalization adaptation. The proposed methods also achieve more than 80% mean accuracy among seafloor, airplane, and ship categories in SSS images.
Keywords:
Generalization recognition
Prompt learning
Mixup technique
Domain adaptation

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

H
Harbin Engineering University
Scholars:
1.9W
Papers: 1.3W
Citations: 1.3W
H
Harbin Medical University
Scholars:
2.9W
Papers: 1.3W
Citations: 1.6W