arrow
Return

Exploiting temporal coherence for self-supervised visual tracking by using vision transformer

delete2022-09-01
delete12
PRE
AI
W
Wenjun Zhu
Z
Zuyi Wang
X
Xu Li *
J
Jun Meng
DOI:10.1016/j.knosys.2022.109318delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning based fully-supervised visual trackers entail the requirement of large-scale and frame -wise annotation that needs a laborious and tedious data annotation process. To reducing the amount of labeled efforts, a self-supervised learning framework, the ETC, is proposed in this work that exploits temporal coherence as a self-supervised signal and uses visual transformer to capture the relationship among the unlabeled video frames. We design a cycle-consistent transformer architecture to cast self -supervised tracking as cycle prediction problems. With carefully-designed and targeted configurations for cycle-consistent transformer including temporal sampling strategies, tracking initialization and data augmentation, our approach is applicable for two tracking settings, i.e., the unlabeled sample (ULS) scene and the few labeled sample (FLS) scene. To learn richer and more discriminative representations, we not only utilize the inter-frame correspondence, but also conduct the intra-frame correspondence to effectively model the target-to-frame and long-range correspondence. Extensive experiments are conducted on the popular benchmark datasets OTB2015, VOT2018, UAV123, TColor-128, NFS and LaSOT, and the results show that our approach achieves competitive results in the ULS setting, and supplies a trade-off between performance and annotation cost in the FLS setting. (C) 2022 Elsevier B.V. All rights reserved.
Keywords:
Self-supervised learning
Visual tracking
Temporal coherence
Transformer

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

Z
zhejiang university
Scholars:
17.7W
Papers: 12.1W
Citations: 152
Cited Papers

Cited Papers

errShare
errSave
Trust in Virtual Teams: A Multidisciplinary Review and Integration
err2019-01-21
err0
errOAAI
errJanine Viol Hacker; Michael Johnson; Carol Saunders; Amanda L. Thayer
errShare
errSave
Unsupervised Deep Representation Learning for Real-Time Tracking
err2020-09-21
err88
errOAAI
errWang, Ning; Zhou, Wengang; Song, Yibing; Ma, Chao; Liu, Wei; Li, Houqiang
errShare
errSave
researcher View more