arrow
Return

Open-World Point Cloud Semantic Segmentation: A Human-in-the-Loop Framework

delete2025-08-06
delete0
PRE
AI
P
Peng Zhang
S
Songru Yang
J
Jinsheng Sun
李蔚清 (Weiqing Li)
Z
Zhiyong Su
DOI:10.1109/TCSVT.2025.3596238delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Open-world point cloud semantic segmentation (OW-Seg) aims to predict point labels for both base and novel classes in real-world scenarios. However, existing methods depend on additional samples to support the prediction on query sample, which is labor-intensive to collect and annotate. Some methods further rely on additional learning stages to adapt to novel classes, limiting their practicality in dynamic scenarios. In addition, the intra-class distribution shifts across samples introduce biased class representations (prototypes), resulting in sub-optimal predictions. To address these limitations, we propose HOW-Seg, the first human-in-the-loop framework for OW-Seg. Instead of relying on additional samples, HOW-Seg constructs class prototypes in the query sample feature space based on sparse point-level annotations, thereby avoiding cross-sample distribution shifts. Considering the lack of granularity of initial prototypes, we introduce an interactive prototype disambiguation mechanism to refine ambiguous prototypes. To further enrich contextual awareness, we propose a prototype label assignment module, which employs a dense conditional random field (CRF) upon the prototypes to optimize their label assignments. Through iterative human feedback, HOW-Seg dynamically improves its predictions, achieving high-quality segmentation for both base and novel classes. Experiments demonstrate that with sparse annotations (e.g., one-class-one-click), HOW-Seg surpasses the state-of-the-art generalized few-shot segmentation (GFS-Seg) method under the 5-shot setting. When using advanced backbones (e.g., Stratified Transformer) and denser annotations (e.g., 10 clicks), HOW-Seg achieves 85.27% mIoU on S3DIS and 66.37% mIoU on ScanNetv2, significantly outperforming other alternatives. The source code will be publicly available at https://github.com/Pengz98/HOW-Seg
Keywords:
Point clouds
semantic segmentation
open-world
human-in-the-loop
prototypes

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

N
Nanjing University of Science and Technology
Scholars:
5.6K
Papers: 2.2K
Citations: 25