返回
Exploiting temporal coherence for self-supervised visual tracking by using vision transformer
DOI:10.1016/j.knosys.2022.109318.png)
摘要
En 中文
Deep learning based fully-supervised visual trackers entail the requirement of large-scale and frame -wise annotation that needs a laborious and tedious data annotation process. To reducing the amount of labeled efforts, a self-supervised learning framework, the ETC, is proposed in this work that exploits temporal coherence as a self-supervised signal and uses visual transformer to capture the relationship among the unlabeled video frames. We design a cycle-consistent transformer architecture to cast self -supervised tracking as cycle prediction problems. With carefully-designed and targeted configurations for cycle-consistent transformer including temporal sampling strategies, tracking initialization and data augmentation, our approach is applicable for two tracking settings, i.e., the unlabeled sample (ULS) scene and the few labeled sample (FLS) scene. To learn richer and more discriminative representations, we not only utilize the inter-frame correspondence, but also conduct the intra-frame correspondence to effectively model the target-to-frame and long-range correspondence. Extensive experiments are conducted on the popular benchmark datasets OTB2015, VOT2018, UAV123, TColor-128, NFS and LaSOT, and the results show that our approach achieves competitive results in the ULS setting, and supplies a trade-off between performance and annotation cost in the FLS setting. (C) 2022 Elsevier B.V. All rights reserved.
Keyword:
Self-supervised learning
Visual tracking
Temporal coherence
Transformer
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
引用论文
Cobalt oxides nanoparticles supported on nitrogen-doped carbon nanotubes as high-efficiency cathode catalysts for microbial fuel cells负载在氮掺杂碳纳米管上的钴氧化物纳米颗粒作为微生物燃料电池的高效阴极催化剂
A novel flake-ball-like magnetic Fe3O4/γ-MnO2 meso-porous nano-composite: Adsorption of fluorinion and effect of water chemistry
Chemosphere
IF0

