arrow
Return

Learning disentangled representation for self-supervised video object segmentation

delete2022-04-01
delete7
PRE
AI
W
Wenjie Hou
Z
Zheyun Qin
X
Xiaoming Xi
X
Xiankai Lu
尹义龙 cover
尹义龙 (Yilong Yin) *
DOI:10.1016/j.neucom.2022.01.066delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper proposes a novel self-supervised method for one-shot video object segmentation where the object annotations are only provided in the first frame. The current self-supervised video object segmentation approaches are implemented by modeling the pairwise correspondence between the target and reference frames. The pairwise correspondence only maintains spatio-temporal consistency. However, the VOS tasks not only require a spatio-temporal relationship between the two frames but also require the salient object information for each frame. In order to achieve this goal, we propose a disentangled representation strategy to disentangle the temporal correspondence into the pairwise term and unary term. The pairwise and unary terms capture inter-frame spatio-temporal and intra-frame salient object information, respectively. To demonstrate the importance of the disentangled representation, we apply the proposed approach to DAVIS-2017 and YouTube-VOS datasets. Experimental results confirm the effectiveness of the proposed solution. (c) 2022 Elsevier B.V. All rights reserved.
Keywords:
Self-supervised video object segmentation
Disentangled representation
Pair-wise term
Unary term

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

S
shandong jianzhu university
Scholars:
4.3K
Papers: 3.1K
Citations: 3
S
shandong university
Scholars:
9.3W
Papers: 6.4W
Citations: 94