arrow
Return

Video object segmentation by multi-scale attention using bidirectional strategy

delete2024-08-01
delete0
PRE
AI
J
Jingxin Wang
Y
Yunfeng Zhang *
F
Fangxun Bao
Y
Yuetong Liu
Q
Qiuyue Zhang
张彩明 (Caiming Zhang)
DOI:10.1016/j.imavis.2024.105136delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper focuses on semi-supervised video object segmentation (VOS). Recently, several Space-Time Memory based networks have effectively improved the performance of VOS. However, most methods predict the target object mask forwardly, which causes error propagation to mislead the future frame segmentation. Moreover, the rich multi-scale information of objects needs to be effectively exploited in videos to extract fine-grained multiscale spatial information. To address these limitations, we present a network with a multi-scale attention module for semi-supervised VOS, which combines a new bidirectional strategy during training. Firstly, we propose the bidirectional strategy in which a backward flow combines the existing standard forward flow. With the strategy, we can rely on the first frame's ground-truth mask to mitigate the problem of error propagation. Secondly, a multi-scale attention module is designed to extracts multi-scale features by different weights and interacts with information between multi-scale channel attention. Especially the multi-scale attention module can effectively extract the fine-grained mask by the network during the bidirectional training. Experimental results show that our network achieves significant segmentation performance compared to state-of-the-art approaches on the YouTube-VOS and DAVIS datasets.
Keywords:
Video object segmentation
Training strategy
Attention mechanism

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

S
shandong university
Scholars:
9.3W
Papers: 6.4W
Citations: 94