arrow
Return

Occluded Video Instance Segmentation: A Benchmark

delete2022-06-18
delete47
delete
OA
AI
J
Jiyang Qi
Y
Yan Gao
Y
Yao Hu
X
Xinggang Wang
X
Xiaoyu Liu
X
Xiang Bai
S
Serge Belongie
A
Alan Yuille
P
Philip H. S. Torr
S
Song Bai *
DOI:10.1007/s11263-022-01629-1delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Can our video understanding systems perceive objects when a heavy occlusion exists in a scene? To answer this question, we collect a large-scale dataset called OVIS for occluded video instance segmentation, that is, to simultaneously detect, segment, and track instances in occluded scenes. OVIS consists of 296k high-quality instance masks from 25 semantic categories, where object occlusions usually occur. While our human vision systems can understand those occluded instances by contextual reasoning and association, our experiments suggest that current video understanding systems cannot. On the OVIS dataset, the highest AP achieved by state-of-the-art algorithms is only 16.3, which reveals that we are still at a nascent stage for understanding objects, instances, and videos in a real-world scenario. We also present a simple plug-and-play module that performs temporal feature calibration to complement missing object cues caused by occlusion. Built upon MaskTrack R-CNN and SipMask, we obtain a remarkable AP improvement on the OVIS dataset. The OVIS dataset and project code are available at http://songbai.site/ovis.
Keywords:
Video instance segmentation
Occlusion reasoning
Dataset
Video understanding
Benchmark

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

A
alibaba group
Scholars:
1.1K
Papers: 789
Citations: 0
U
University of Copenhagen
Scholars:
7.6W
Papers: 6.6W
Citations: 86
J
Johns Hopkins University
Scholars:
10.2W
Papers: 8.8W
Citations: 13.0W
U
university of oxford
Scholars:
9.6W
Papers: 8.5W
Citations: 137
researcher View more organizations