arrow
Return

Self-supervised vision transformers for semantic segmentation

delete2025-02-01
delete0
PRE
AI
X
Xianfan Gu
Y
Yingdong Hu
C
Chuan Wen
Y
Yang Gao *
DOI:10.1016/j.cviu.2024.104272delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Semantic segmentation is a fundamental task in computer vision and it is a building block of many other vision applications. Nevertheless, semantic segmentation annotations are extremely expensive to collect, sousing pre-training to alleviate the need fora large number of labeled samples is appealing. Recently, self-supervised learning (SSL) has shown effectiveness in extracting strong representations and has been widely applied to a variety of downstream tasks. However, most works perform sub-optimally in semantic segmentation because they ignore the specific properties of segmentation: (i) the need of pixel level fine-grained understanding; (ii) with the assistance of global context understanding; (iii) both of the above achieve with the dense self- supervisory signal. Based on these key factors, we introduce a systematic self-supervised pre-training framework for semantic segmentation, which consists of a hierarchical encoder-decoder architecture MEVT for generating high-resolution features with global contextual information propagation and a self-supervised training strategy for learning fine-grained semantic features. In our study, our framework shows competitive performance compared with other main self-supervised pre-training methods for semantic segmentation on COCO-Stuff, ADE20K, PASCAL VOC, and Cityscapes datasets. e.g., MEVT achieves the advantage in linear probing by +1.3 mIoU on PASCAL VOC.
Keywords:
Self-supervised representation learning
Semantic segmentation
Vision transformer

Journal

Computer Vision and Image Understanding cover
Computer Vision and Image Understanding
IF:
3.5
Papers:
428
Citations:
7.3K

Organization

S
shanghai qi zhi institute
Scholars:
64
Papers: 45
Citations: 0