Return
DFCTrack: dynamic full-temporal context fusion for object tracking
DOI:10.1007/s13042-026-03261-8.png)
Abstract
En 中文
Utilizing rich video-level contextual information to perceive targets effectively is crucial for object tracking. However, existing methods typically capture limited-length context through simple multiple video or adjacent frames, resulting in underutilization of contextual information. To address this issue, we propose the DFCTrack tracking framework. It captures target behavior changes across the entire temporal span of the complete video sequence rather than a few images. Specifically, we propose a long-term context modeling mechanism based on collaborative architecture between Mamba and backbone networks. It utilizes Mamba’s hidden state iteration to enhance the backbone’s cross-frame spatiotemporal information aggregation capability. This integrates target evolution cues across full temporal dimensions and captures long-range spatiotemporal information. On the other hand, we propose a Short-term Target Redundancy Removal module. This module precisely selects high-value tokens relevant to the current target by jointly optimizing attention entropy and classification score. Provides recent high-quality target representations for the tracker. Finally, we design a Long-Short Term Context Fusion module, which leverages a self-attention mechanism to dynamically weight and fuse long-term context with recent high-quality token representations. This introduces global temporal information into the model, significantly enhancing its spatiotemporal awareness. Experiments demonstrate that DFCTrack achieves excellent tracking performance on multiple benchmarks such as GOT-10K, TrackingNet, and LaSOT.
Keywords:
Object tracking
Long-term context modeling
Spatiotemporal representation
Journal
IF:
2.7
Papers:
3.1K
Citations:
5.6K

