返回
Multi-view consistent feature learning for open-set semantic image segmentation
DOI:10.1016/j.eswa.2025.130791.png)
摘要
En 中文
Recently, open-set semantic segmentation has attracted increasing attention in computer vision. However, existing methods are generally trained by utilizing a set of images with limited viewpoints, so that they could not effectively segment various partially occluded objects of unseen classes. To address this problem, we propose a Multi-view Consistent Feature Learning method for open-set semantic image segmentation, which called MCFL. The proposed MCFL consists of a view-consistent feature learning network and an open-set segmentation block. The view-consistent feature learning network employs neural radiance fields for feature extraction and generates a continuous 3D spatial representation, which furnishes a view-consistent feature description. Moreover, it aligns features of seen classes closely with their prototypes in the metric space, enhancing the network's ability to maintain class coherence in its representations. The open-set segmentation block comprises an unseen-class prediction unit and an open-set segmentation map prediction unit. In the unseen-class prediction unit, we employ a hybrid distance sum criterion to calculate the unseen-class probability for each pixel. Subsequently, in the open-set segmentation map prediction unit, different labels are assigned to each pixel based on the probability of belonging to the unseen-class and the distribution of seen-class features, thereby generating the open-set segmentation map. Experimental results on Replica and ScanNet datasets demonstrate that the proposed MCFL outperforms the comparative methods in most cases. The implementation is publicly available at https://github.com/bowen099/MCFL.
Keyword:
Open-set semantic segmentation
View-consistent feature learning
Neural radiance fields
Partially occluded
3D scene understanding
Metric learning
Hybrid distance sum criterion
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W

