Return
LoongTrack: Exploring long-sequence modeling for visual tracking
DOI:10.1016/j.neunet.2026.108824.png)
Abstract
En 中文
Visual object tracking constitutes a complex online association problem, characterized by its intricate interplay of temporal and spatial dimensions. In the current mainstream pipeline, images are typically encoded into a sequence, which is subsequently processed and fused by a transformer. However, due to the quadratic complexity of transformers, they can only carefully design complex fusion modules to extract limited spatiotemporal cues, while neglecting the spatiotemporal information inherently provided by the video frames themselves. In this paper, rather than meticulously designing network modules, we explore a foundational tracker with linear complexity to exploit the information inherent in the frames. Specifically, we improve the video scanning approach by proposing causal consistent scanning, which accounts for the causal nature of tracking sequences. Based on this, we build a simple tracking pipeline to explore long sequence modeling (i.e., long time series, and large resolution) for capturing the successive information flow under dense contexts. Moreover, we systematically develop selective scan patterns from the dual perspectives of temporal and spatial dimensions in long sequences and attempt to conduct a detailed analysis. Extensive experiments on multiple public datasets demonstrate that our approach, which employs fewer parameters and training memory consumption, achieves encouraging results. We hope that this work can provide valuable inspiration for enriching spatiotemporal information in tracking task, and the code will be made available.
Keywords:
Visual object tracking
Spatiotemporal modeling
Long sequence processing
Causal consistent scanning
Linear complexity
Journal
IF:
6.3
Papers:
7.8K
Citations:
3.0W

