Return
Class-Guided Visual-Language Prompt Learning Model for Remote Sensing Image Classification
DOI:10.1109/lgrs.2026.3708468.png)
Abstract
En 中文
Vision-language models (VLMs) like contrastive language-image pre-training (CLIP), renowned for powerful zero-shot capabilities, offer a promising approach to mitigating data scarcity and limited generalization in remote sensing (RS). Prompt learning offers parameter-efficient adaptation, but existing methods typically rely on static text prompts and global visual features, limiting generalization to unseen classes. To address these issues, first, a class-guided prompt learning (CGPL) text-encoder is designed; through a text class information embedding module (TCIEM), class-level semantic priors are injected into learnable prompts to generate class-aware dynamic prompts, overcoming static text prompts limitations. Second, a hierarchical multiscale feature fusion (HMSFF) image encoder is proposed, which injects multiscale features into the global representation through adaptive adapter feature fusion (AAFF) module, improving the model’s multiscale representation learning ability. Finally, a multitask contrastive learning (MTCL) loss is constructed to balance the classification and generalization capabilities. Extensive experiments on four public RS datasets demonstrate our method achieves superior average performance. Notably, it significantly outperforms state-of-the-art methods on complex datasets characterized by multiscale objects and dense backgrounds (e.g., RSICD, RESISC45, MLRSNet). While remaining competitive on simpler macro-pattern datasets (e.g., PatternNet), our approach exhibits robust representation learning for seen classes and strong zero-shot generalization for unseen ones.
Keywords:
Contrastive language-image pre-training (CLIP)
multiscale feature fusion
prompt learning
remote sensing (RS) image scene classification
Journal
I
IF:
4.4
Papers:
560
Citations:
0

