arrow
Return

From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature Mining

delete2022-01-01
delete4
PRE
AI
R
Ruoqi Li
王一帆 cover
王一帆 (Yifan Wang) *
L
Lijun Wang
卢湖川 (Huchuan Lu)
X
Xiaopeng Wei
张强 (Qiang Zhang)
DOI:10.1109/TIP.2022.3201603delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Existing self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts.
Keywords:
Semantics
Training
Feature extraction
Image reconstruction
Task analysis
Object segmentation
Image segmentation
Video object segmentation
self-supervised learning
pixel-level correspondence
semantic-level adaption
feature mining

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

D
Dalian University of Technology
Scholars:
5.9W
Papers: 4.4W
Citations: 5.5W