Return
Unlocking the power of vision foundation models via multi-expert collaboration for cross-domain few-shot segmentation
L
Z
C
Z
DOI:10.1016/j.knosys.2026.116039.png)
Abstract
En 中文
Although cross-domain few-shot segmentation (CD-FSS) has achieved promising results in segmenting novel classes across domains with limited samples, most existing approaches predominantly focus on enhancing domain adaptability through simple backbone architectures. Few studies have explored the potential of using vision foundation models (VFMs), which present a compelling solution for cross-domain tasks owing to their exceptional generalization capabilities. Nevertheless, we find that directly applying VFMs (e.g., DINOv2) to CD-FSS still faces three major challenges: (1) direct fine-tuning risks compromising the model's inherent representation capacity; (2) features at different stages demonstrate distinct functional characteristics, requiring careful utilization strategies; (3) the attention region may shift. To this end, we propose three innovative solutions. First, we freeze the VFM and introduce dual adaptation modules at the image and frequency levels to effectively enhance prototype matching and cross-domain adaptability. Second, our multi-expert collaborative optimization strategy leverages the complementary advantages of multi-stage features through parallel prediction branches and a learnable domain-specific fusion strategy, effectively solving the hierarchical feature optimization imbalance problem during source domain training. Finally, we propose the text-guided prompt correction module, which employs text embeddings to refine visual features and improve their focus on target regions. Extensive experiments show that our method significantly outperforms state-of-the-art CD-FSS methods.
Keywords:
Cross-domain
Few-shot
Semantic segmentation
Vision foundation models
Multi-expert collaboration
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
