返回
DiVOT: Differentiated Interaction-Guided Video-Level Object Tracking
DOI:10.1016/j.dsp.2026.105955.png)
摘要
En 中文
Recent advancements in video-level methods have made significant strides in the object tracking field. This method leverages multiple online templates to capture rich temporal information. However, most existing methods treat online templates as equally important as the initial template, overlooking the inherent instability of online templates during updating, which consequently degrades tracking performance. To alleviate this issue, we propose a novel differentiated interaction-guided video-level object tracking method, termed DiVOT, aimed at mitigating the impact of template instability and boosting the tracking performance. Our feature extraction network consists of a differentiated encoder block, which differentially guides the interaction between the search region and various templates, enabling the tracker to achieve a balance between stability and adaptability. Additionally, we design an auxiliary module, i.e., the memory decoder, to compensate for the deficiency of the differentiated interaction, where the latency of online templates hinders the acquisition of the most recent target appearance information. Extensive experiments on six mainstream datasets, i.e., OTB100, GOT-10k, TrackingNet, VOT2020, NFS, and LaSOT, validate the effectiveness of our proposed method.
期刊
D
IF:
3
论文数:
768
被引数:
0
机构
暂无机构信息
引用论文
Occlusion-aware visual object tracking based on multi-template updating Siamese network基于多模板更新Siamese网络的遮挡感知视觉目标跟踪
EMAT: Efficient feature fusion network for visual tracking via optimized multi-head attentionEMAT: 通过优化的多头注意力进行视觉跟踪的高效特征融合网络
NEURAL NETWORKS
IF6.3
Learning orientational interaction-aware attention and localization refinement for object tracking学习具有方向交互感知的注意力和定位精修用于目标跟踪
Efficient long-term tracking with local-global similar object interference suppression高效长期跟踪与局部-全局相似目标干扰抑制

