arrow
Return

Hierarchical Open-vocabulary Part-object Segmentation with Knowledge-guided SAM

delete2026-02-02
delete0
PRE
AI
X
Xin-Jian Wu
刘程琳 (Cheng‐Lin Liu) *
DOI:10.1007/s11633-025-1610-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Open-vocabulary semantic segmentation (OV-Seg) aims to segment novel categories without dense annotations, facilitating scalable perception in open-world scenarios. While recent advances in foundation models–such as CLIP for semantic grounding and segment anything model (SAM) for spatial localization–have boosted OV-Seg performance, most existing methods remain limited to object-level segmentation and neglect the compositional and hierarchical structure of real-world entities. In this paper, we propose a hierarchical open-vocabulary segmentation framework that integrates CLIP and SAM with structured external knowledge in the form of object–part hierarchies. Specifically, we construct a category-level knowledge graph to encode part–whole relationships and guide the generation of enriched prompts, thereby aligning visual features with both object-level and part-level semantics. Extensive experiments on PartImageNet and PASCAL-Part demonstrate that our method consistently outperforms state-of-the-art baselines, especially in part-level segmentation and novel-category generalization. These results confirm the effectiveness of incorporating structured priors to enhance compositional and fine-grained visual understanding in open-vocabulary settings.
Keywords:
Open-vocabulary segmentation
fine-grained segmentation
vision-language foundation models
knowledge-guided prompting
object–part hierarchy

Journal

Machine Intelligence Research cover
Machine Intelligence Research
IF:
8.7
Papers:
301
Citations:
882

Organization

I