arrow
返回

Transformer-based Cross Reference Network for video salient object detection

delete2022-08-01
delete32
PRE
AI
K
Kan Huang
C
Chunwei Tian *
苏
苏敬勇 (Jingyong Su)
J
Jerry Chun‐Wei Lin
DOI:10.1016/j.patrec.2022.06.006delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video salient object detection is a fundamental computer vision task aimed at highlighting the most conspicuous objects in a video sequence. There are two key challenges presented in video salient ob-ject detection: (1) how to extract effective feature representations from appearance and motion cues, and (2) how to combine both of them into robust saliency representation. To handle these challenges, in this paper, we propose a novel Transformer-based Cross Reference Network (TCRN), which fully exploits long-range context dependencies in both feature representation extraction and cross-modal (i.e., appear-ance and motion) integration. In contrast to existing CNN-based methods, our approach formulates video salient object detection as a sequence-to-sequence prediction task. In the proposed approach, the deep feature extraction is achieved by a pure vision transformer with multi-resolution token representations. Specifically, we design a Gated Cross Reference (GCR) module to effectively integrate appearance and motion into saliency representation. The GCR first propagates global context information between differ-ent modalities, and then perform cross-modal fusion by a gate mechanism. Extensive evaluations on five widely-used benchmarks show that the proposed Transformer-based method performs favorably against the existing state-of-the-art methods (c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Video salient
Object detection
Transformer
Cross -modal integration

期刊

Pattern Recognition Letters 封面图
Pattern Recognition Letters
IF:
3.3
论文数:
8.0K
被引数:
1.6W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
W
Western Norway University of Applied Sciences
学者数:
2.2K
论文数: 2.2K
被引数: 1.4K
S
Shanghai Maritime University
学者数:
4.8K
论文数: 4.2K
被引数: 4.7K
N
Northwestern Polytechnical University
学者数:
4.6W
论文数: 3.7W
被引数: 5.3W
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Exploring Rich and Efficient Spatial Temporal Interactions for Real-Time Video Salient Object Detection
err2021-01-01
err84
errOAAI
errChen, Chenglizhao; Wang, Guotao; Peng, Chong; Fang, Yuming; Zhang, Dingwen; Qin, Hong
err分享
err收藏
A Laboratory Model of Urban Street-Canyon Flows
err2000-09-01
err0
errOAAI
errJong-Jin Baik; Rae-Seol Park; Hye-Yeong Chun; Jae-Jin Kim
err分享
err收藏
err分享
err收藏
SCOM: Spatiotemporal Constrained Optimization for Salient Object Detection
err2018-07-01
err92
PREAI
errChen, Yuhuan; Zou, Wenbin; Tang, Yi; Li, Xia; Xu, Chen; Komodakis, Nikos
err分享
err收藏
Radiation‐induced sensitivity of tissue‐resident mesenchymal stem cells in the head and neck region
err2019-04-24
err0
PREAI
errJennifer L. Spiegel; Mario Hambrecht; Vera Kohlbauer; Frank Haubner; Friedrich Ihler; Martin Canis; Arndt F. Schilling; Kai O. Böker; Ralf Dressel; Katrin Streckfuss‐Bömeke; Mark Jakob
err分享
err收藏
Video Saliency Detection via Spatial-Temporal Fusion and Low-Rank Coherency Diffusion
err2017-07-01
err159
errOAAI
errChen, Chenglizhao; Li, Shuai; Wang, Yongguang; Qin, Hong; Hao, Aimin
err分享
err收藏
学者 查看更多内容