Return
MSPD-net: structural–appearance prototype decoupling for weakly supervised semantic segmentation
DOI:10.1007/s00371-026-04715-4.png)
Abstract
En 中文
Weakly supervised semantic segmentation (WSSS) based on image-level labels typically relies on class activation maps (CAMs) to generate pseudo-labels. However, single-layer prototype-based generation methods often result in incomplete CAM activation and inaccurate target localization due to limited feature representation capabilities. Furthermore, the reliance of existing methods on discriminative appearance information makes it difficult to effectively utilize object structural information, further impacting boundary accuracy and the final segmentation performance. To address this problem, this paper proposes a complementary prototype representation framework, employing three modules to collaboratively improve pseudo-label quality. First, the Multi-scale prototype fusion module (MSPF) adaptively integrates multi-level features to construct a more complete and accurate category representation. Second, the structural–appearance prototype decoupling module (SAPD) separates appearance and geometric information in pixel-prototype matching, enhancing texture discrimination and boundary details. Finally, Negative prototype clustering mining (NPCM) improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples. Experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate that our method achieves 73.3% mIoU (val) and 73.7% mIoU (test) on PASCAL VOC, and 43.7% mIoU on MS COCO. The code is available at https://github.com/Weiw1819/MSPD-Net .
Keywords:
Weakly supervised semantic segmentation
Class activation map
Prototype learning
Structural–appearance decoupling
Multi-scale feature fusion
Contrastive learning
Boundary-aware segmentation
Journal
IF:
2.9
Papers:
4.6K
Citations:
6.5K
Organization
Cited Papers
Context-aware learning and background activation suppression for weakly supervised semantic segmentation
MULTIMEDIA SYSTEMS
IF3.1
Weakly supervised semantic segmentation via saliency perception with uncertainty-guided noise suppression
VISUAL COMPUTER
IF2.9
no more

