Return
DU3D: Multi-View Denoising for Unsupervised 3D Instance Segmentation
DOI:10.1109/tcsvt.2026.3735777.png)
Abstract
En 中文
The rapid progress of vision foundation models (VFMs) has advanced unsupervised 3D instance segmentation, where VFM features are lifted into 3D to generate pseudo masks. However, lifted features often contain noisy artifacts, forcing existing methods to rely on additional 2D or 3D unsupervised models for refinement. Common denoising strategies mainly use feature smoothing, which blurs distinct representations, weakens discriminability, and degrades pseudo-mask quality. To address this limitation, we propose DU3D, whose training-free pseudo-mask generation stage denoises lifted visual features via cross-view consistency and generates pseudo masks using only denoised features and 3D geometric priors. DU3D consolidates cross-view VFM feature consistency and aggregates multi-view features to construct a noise prototype, which guides the formation of a denoised 3D feature field. We then observe that this field exhibits strong part-level priors. Leveraging this property, we introduce a simple yet effective Part-Aware Clustering algorithm that fully unleashes the potential of VFM features to produce high-quality pseudo masks, eliminating the need for extra unsupervised models. Finally, we obtain the final segmentation via denoising self-training with a robust loss that mitigates pseudo-label noise. Extensive experiments on ScanNetV2, ScanNet++, and S3DIS demonstrate DU3D’s superiority over existing methods.
Keywords:
3D Instance segmentation
Unsupervised learning
Indoor scene
Journal
IF:
11.1
Papers:
846
Citations:
3.1W
Organization
Cited Papers
No cited papers available

