Return
FocTrack: Focus attention for visual tracking
DOI:10.1016/j.patcog.2024.111128.png)
Abstract
En 中文
Transformer trackers have achieved widespread success based on their attention mechanism. The vanilla attention mechanism focuses on modeling the long-range dependencies between tokens to gain a global perspective. However, inhuman tracking behavior, the line of sight first skims apparent regions and then focuses on the differences between similar regions. To explore this issue, we build a powerful online tacker with focus attention, named FocTrack. Firstly, we design a focus attention module, which adopts the iterative binary clustering function (IBCF) before self-attention to simulate human behavior. Specifically, fora given cluster, other clusters are treated as apparent tokens that are skimmed during the clustering process, while the subsequent self-attention performs focused discriminative learning on the target cluster. Moreover, we propose a local template update strategy (LTUS) to probe into the effective temporal information for visual object tracking. In the testing, LTUS only replaces outdated local templates to ensure overall reliability and holds a low computational burden. Finally, extensive experiments show that our proposed FocTrack achieves state-of-the-art performance in several benchmarks.In particular, FocTrack achieves 71.5% AUC on the LaSOT, 84.7% AUC on the TrackingNet, and a running speed of around 36 FPS, outperforming the popular approaches.
Keywords:
Visual object tracking
Focus attention
Clustering function
Update strategy
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

