Return
DiVOT: Differentiated Interaction-Guided Video-Level Object Tracking
DOI:10.1016/j.dsp.2026.105955.png)
Abstract
En 中文
Recent advancements in video-level methods have made significant strides in the object tracking field. This method leverages multiple online templates to capture rich temporal information. However, most existing methods treat online templates as equally important as the initial template, overlooking the inherent instability of online templates during updating, which consequently degrades tracking performance. To alleviate this issue, we propose a novel differentiated interaction-guided video-level object tracking method, termed DiVOT, aimed at mitigating the impact of template instability and boosting the tracking performance. Our feature extraction network consists of a differentiated encoder block, which differentially guides the interaction between the search region and various templates, enabling the tracker to achieve a balance between stability and adaptability. Additionally, we design an auxiliary module, i.e., the memory decoder, to compensate for the deficiency of the differentiated interaction, where the latency of online templates hinders the acquisition of the most recent target appearance information. Extensive experiments on six mainstream datasets, i.e., OTB100, GOT-10k, TrackingNet, VOT2020, NFS, and LaSOT, validate the effectiveness of our proposed method.
Journal
D
IF:
3
Papers:
653
Citations:
0
Organization
No organization information available

