Return
Adaptive representation-aligned modeling for visual tracking
DOI:10.1016/j.knosys.2024.112847.png)
Abstract
En 中文
Accurate object representations and reliable object states are essential for robust visual tracking. While current Transformer-based trackers employ symmetric and asymmetric attention mechanisms to learn object representations, they often overlook attention discrepancies caused by the target object and distractors between the template and search region. These discrepancies can lead to the inclusion of distractors, which degrade the quality of the learned representations. To address this issue, we propose ARATrack, an adaptive representation- aligned tracker. Specifically, we present a representation-aligned attention (RAA) mechanism, which adaptively refines object representations through a novel attention alignment strategy. This strategy ensures consistent attention between the template and search region, while minimizing interference from distractors. Furthermore, to handle dynamic changes in object states, we propose a state-aware head. This module continuously updates the object state using the refined representations, enhancing the tracker's resilience to appearance changes and improving its stability in diverse and challenging conditions. Together, these components work synergistically, offering mutual guidance to ensure that ARATrack delivers stable, reliable, and high-performance visual tracking. Extensive experiments on six benchmark datasets demonstrate that ARATrack outperforms other state-of-the-art trackers, achieving competitive performance across a diverse range of tracking scenarios. Our code is publicly available at https://github.com/nubsym/ARATrack.
Keywords:
Visual object tracking
Transformer trackers
Representation-aligned attention
Adaptive state-aware
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

