Return
Image-text driven style randomization for domain generalized semantic segmentation
DOI:10.1016/j.neucom.2026.132953.png)
Abstract
En 中文
Semantic segmentation models trained on source domains often fail to generalize to unseen domains due to domain shifts caused by varying environmental conditions. While existing approaches rely solely on text prompts for domain randomization, their generated styles often deviate from real-world distributions. To address this limitation, we propose a novel two-stage framework for Domain Generalization in Semantic Segmentation (DGSS). First, we introduce Image-Prompt-driven Instance Normalization (I-PIN), which leverages both style images and text prompts to optimize style parameters, achieving more accurate style representations compared to text-only approaches. Second, we present Dual-Path Style-Invariant Feature Learning (DSFL) that employs inter-style and intra-style consistency losses, ensuring consistent predictions across different styles while promoting feature alignment within semantic classes. Extensive experiments demonstrate that our approach consistently outperforms existing state-of-the-art methods across multiple challenging domains, effectively addressing the domain shift problem in semantic segmentation.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

