arrow
Return

CostDiff: Residual Diffusion-Based Cost Map Refinement for Open-Vocabulary Semantic Segmentation

delete2026-01-01
delete0
PRE
AI
B
Bowen Deng
Y
Yutao Rao
F
Fangyu Wu
J
Junjie Zhang *
DOI:10.1007/978-981-95-5761-5_9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Open-Vocabulary Semantic Segmentation (OVSS) empowers models to recognize novel classes beyond predefined categories. While contrastive Vision-Language Models (VLMs) like CLIP enable open-vocabulary learning, they struggle with pixel-level semantic localization due to image-level pretraining. We propose a residual diffusion-based cost map refinement strategy to address these challenges. By treating CLIP's coarse-grained classification maps as initial cost maps, our method iteratively refines them via a multi-step diffusion process, bridging the gap between high-level semantics and low-level spatial details. This enhances pixel-wise discriminative ability without retraining VLMs. Experiments on standard benchmarks demonstrate promising improvements in both quantitative accuracy and qualitative boundary precision, verifying the effectiveness of integrating diffusion for OVSS. Our approach offers a novel paradigm for advancing open-vocabulary visual understanding via foundation model refinement.
Keywords:
Open-Vocabulary Semantic Segmentation
Residual Diffusion
Cost Map Refinement

Journal

P
PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2025, PT XII
IF:
0
Papers:
27
Citations:
0

Organization

X
xi'an jiaotong-liverpool university
Scholars:
1.0K
Papers: 536
Citations: 0
S
shanghai university
Scholars:
3.9W
Papers: 2.7W
Citations: 52