arrow
Return

Real-time human-centric segmentation for complex video scenes

delete2022-10-01
delete1
delete
OA
AI
R
Ran Yu
C
Chenyu Tian
W
Weihao Xia
X
Xinyuan Zhao
L
Liejun Wang
杨余久 (Yujiu Yang) *
DOI:10.1016/j.imavis.2022.104552delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Most existing video tasks related to human focus on the segmentation of salient humans, ignoring the unspec-ified others in the video. Few studies have focused on segmenting and tracking all humans in a complex video, including pedestrians and humans of other states (e.g., seated, riding, or occluded). In this paper, we propose a novel framework, abbreviated as HVISNet, that segments and tracks all presented people in given videos based on a one-stage detector. To better evaluate complex scenes, we offer a new benchmark called HVIS (Human Video Instance Segmentation), which comprises 1447 human instance masks in 805 high-resolution videos in di-verse scenes. Extensive experiments show that our proposed HVISNet outperforms the state-of-the-art methods in terms of accuracy at a real-time inference speed (30 FPS), especially on complex video scenes. We also notice that using the center of the bounding box to distinguish different individuals severely deteriorates the segmen-tation accuracy, especially in heavily occluded conditions. This common phenomenon is referred to as the ambig-uous positive samples problem. To alleviate this problem, we propose a mechanism named Inner Center Sampling to improve the accuracy of instance segmentation. Such a plug-and-play inner center sampling mech-anism can be incorporated in any instance segmentation model based on a one-stage detector to improve the performance. In particular, it gains 4.1 mAP improvement on the state-of-the-art method in the case of occluded humans. Code and data are available at https://github.com/IIGROUP/HVISNet.(c) 2022 Elsevier B.V. All rights reserved.
Keywords:
Multiple human tracking
Video instance segmentation
One -stage detector
Video understanding
Deep neural networks
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

H
huawei technologies
Scholars:
3.3K
Papers: 2.9K
Citations: 1
T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137
T
Tsinghua Shenzhen International Graduate School
Scholars:
6.8K
Papers: 4.9K
Citations: 9
researcher View more organizations