Return
Multi-modal recursive prompt learning with mixup embedding for generalization recognition
DOI:10.1016/j.knosys.2024.111726.png)
Abstract
En 中文
The contrastive language-image pretraining (CLIP) model has shown promise in generalization recognition by combining visual and textual embeddings. However, the pretrained CLIP model requires further fine tuning for downstream tasks. Existing prompt learning (PL) approaches, while effective, neglect the cross-hierarchical fusion of multimodal features. To address this issue, we introduce multi -modal recursive PL (MmRPL) for different generalization image recognition tasks. Recursive connections between prompts at different layers facilitate hierarchy-aware vision -text fusion. To our best knowledge, we introduce a mixup embedding technique to enhance feature representations of PL for the first time. Extensive experiments across popular generalization recognition settings reveal that MmRPL outperforms existing methods. Furthermore, when combined with the embedding technique, MmRPL further enhances recognition in base-to-novel generalization, generalized zeroshot, domain generalization, and domain adaptive zero-shot scenarios. Notably, the proposed methods extend to applications like zero-shot recognition of side-scan sonar (SSS) image targets via domain generalization adaptation. The proposed methods also achieve more than 80% mean accuracy among seafloor, airplane, and ship categories in SSS images.
Keywords:
Generalization recognition
Prompt learning
Mixup technique
Domain adaptation
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

