arrow
返回

Learning Dynamical Position Embedding for Discriminative Segmentation Tracking

delete2024-08-01
delete0
PRE
AI
Y
Yijin Yang
X
Xiaodong Gu *
DOI:10.1109/TITS.2024.3350673delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual tracking plays a pivotal role in intelligent transportation systems and has a wide range of practical applications such as autonomous driving and traffic counting. Recently, the attention mechanism in Transformers has been successfully applied to the field of visual tracking, leading to a significant improvement in tracking performance. However, Transformer-based trackers directly flatten two-dimensional image features into one-dimensional vectors to compute attention scores. This process unavoidably results in the omission of crucial position distribution information necessary for precise target localization. To address this issue, we propose a novel cross-attention based tracking-by-segmentation framework, called Dynamical Position Embedding based Tracking framework (DPET). DPET incorporates an additional network for modeling position information to complement the cross-attention module. To be specific, a dynamical position embedding network is introduced to adaptively encode position information. This network is then integrated into the cross-attention based feature fusion network to compensate for the loss of position distribution information. As a result, the fused feature incorporates abundant contextual semantic cues for target classification and precise position information for target localization simultaneously. To overcome the constraints imposed by bounding-boxes, a segmentation network that takes the fused feature as input is designed to achieve accurate pixel-wise tracking. Extensive experiments on eight challenging tracking benchmarks show that our DPET tracker enables real-time operations and achieves promising tracking performance on the GOT-10K benchmark. Especially, DPET tracker achieves the top accuracy scores on VOT2016, VOT2018 and VOT2019 benchmarks.
Keyword:
Target tracking
Correlation
Feature extraction
Transformers
Convolution
Semantics
Benchmark testing
Visual object tracking
attention mechanism
tracking-by-segmentation
dynamical position embedding

期刊

IEEE Transactions on Intelligent Transportation Systems 封面图
IEEE Transactions on Intelligent Transportation Systems
IF:
8.4
论文数:
9.7K
被引数:
6.3W

机构

F
fudan university
学者数:
11.8W
论文数: 7.7W
被引数: 121
引用论文

引用论文

err分享
err收藏
Trust in Virtual Teams: A Multidisciplinary Review and Integration
err2019-01-21
err0
errOAAI
errJanine Viol Hacker; Michael Johnson; Carol Saunders; Amanda L. Thayer
err分享
err收藏
err分享
err收藏
Evaluating Choroidal Characteristics in Systemic Sclerosis Using Enhanced Depth Imaging Optical Coherence Tomography
err2016-02-22
err0
PREAI
errEbru Esen; Didem Arslan Tas; Selcuk Sizmaz; Ipek Turk; Ilker Unal; Nihal Demircan
err分享
err收藏
Explosive synchronization: From synthetic to real-world networks
err2022-01-01
err0
PREAI
errAtiyeh Bayani; Sajad Jafari; Hamed Azarnoush
err分享
err收藏
学者 查看更多内容