arrow
Return

Ref-Diff: zero-shot referring image segmentation with generative models

delete2026-07-31
delete0
PRE
AI
M
Minheng Ni
Y
Yabo Zhang
K
Kailai Feng
X
Xiaoming Li
Y
Yiwen Guo
左旺孟 (Wangmeng Zuo) *
DOI:10.1007/s11432-024-4985-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Zero-shot referring image segmentation (RIS) is a challenging task that involves identifying an instance segmentation mask from referring texts, without being trained on paired image-text data. Current zero-shot RIS methods mainly rely on pre-trained discriminative models (e.g., CLIP). In contrast, this study investigates the potential of generative models (e.g., Stable Diffusion) to understand relationships between various visual elements and text descriptions, an area that remains unexplored in this context. In this work, we introduce the Referring Diffusional segmentor (Ref-Diff), a model that harnesses the fine-grained multimodal information provided by generative models. Our results demonstrate that Ref-Diff, using only a generative model and no external proposal generator, outperforms state-of-the-art weakly supervised models on the RefCOCO+ and RefCOCOg benchmarks. Furthermore, by combining both generative and discriminative models, we present the enhanced version, Ref-Diff+, which significantly surpasses existing methods. This emphasizes the benefits of generative models for discriminative models, thereby improving referring segmentation.
Keywords:
deep learning
generative model
referring image segmentation
zero-shot learning
zero-shot referring image segmentation

Journal

S
Science China-Information Sciences
IF:
7.6
Papers:
87
Citations:
0

Organization

F
Faculty of Computing
Scholars:
71
Papers: 30
Citations: 0