arrow
Return

DiVOT: Differentiated Interaction-Guided Video-Level Object Tracking

delete2026-01-27
delete0
PRE
AI
Z
Zhixi Wu
陈思 (Si Chen) *
D
Da-Han Wang
S
Shunzhi Zhu *
DOI:10.1016/j.dsp.2026.105955delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advancements in video-level methods have made significant strides in the object tracking field. This method leverages multiple online templates to capture rich temporal information. However, most existing methods treat online templates as equally important as the initial template, overlooking the inherent instability of online templates during updating, which consequently degrades tracking performance. To alleviate this issue, we propose a novel differentiated interaction-guided video-level object tracking method, termed DiVOT, aimed at mitigating the impact of template instability and boosting the tracking performance. Our feature extraction network consists of a differentiated encoder block, which differentially guides the interaction between the search region and various templates, enabling the tracker to achieve a balance between stability and adaptability. Additionally, we design an auxiliary module, i.e., the memory decoder, to compensate for the deficiency of the differentiated interaction, where the latency of online templates hinders the acquisition of the most recent target appearance information. Extensive experiments on six mainstream datasets, i.e., OTB100, GOT-10k, TrackingNet, VOT2020, NFS, and LaSOT, validate the effectiveness of our proposed method.

Journal

D
Digital Signal Processing
IF:
3
Papers:
653
Citations:
0

Organization

No organization information available