Return
TRL: Transformer based refinement learning for hybrid-supervised semantic segmentation
DOI:10.1016/j.patrec.2022.11.015.png)
Abstract
En 中文
This paper studies a new yet practical setting of semi-supervised semantic segmentation, i.e., hybrid-supervised semantic segmentation, where a small number of pixel-level (strong) annotations and a large number of image-level (weak) annotations are provided. It is a common practice to utilize pseudo labels to mitigate the issue of lacking strong annotations. However, most of the existing works focus on improving the model representation with unlabeled data, while ignoring the quality of pseudo labels, leading to poor segmentation performance. It is difficult to directly learn a model with limited images to produce high-quality pseudo labels. To address this problem, we propose a novel learning method, i.e., Transformer based Refinement Learning (TRL), which explores a learning process under the assistance of weak annotations and the supervision of strong annotations. TRL progressively refines heat maps from the poor qualities to the better ones to obtain satisfactory pseudo labels. Specifically, we propose a Dual-Cross Transformer Network (DCTN) to perform the refinement learning. DCTN extracts the features from both images and heat maps by a dual-stream network. The cross attentions inside DCTN hierarchically fuse the dual-stream features. The experiments on the PASCAL VOC and COCO datasets show that TRL outperforms the state-of-the-art methods for hybrid-supervised semantic segmentation. (c) 2022 Elsevier B.V. All rights reserved.
Keywords:
Hybrid-supervised semantic segmentation
Simi-supervised semantic segmentation
Weakly-supervised semantic segmentation
Refinement learning
Heat map
Pseudo label
Journal
IF:
3.3
Papers:
7.9K
Citations:
1.6W

