arrow
Return

Memory optimization network for enhanced vision-language target tracking

delete2026-03-04
delete0
PRE
AI
J
Jianwei Zhang
W
Wendi Zhang
H
Huanlong Zhang
S
Shoukang Yi
B
Bin Jiang
Z
Zhoujingzi Qiu *
DOI:10.1007/s11227-026-08377-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In video-language tracking, providing comprehensive and precise historical information is crucial to handle scenarios involving target appearance variations. However, the historical information utilized by existing methods contains substantial noise and typically requires repetitive training. It leads to imprecise historical reference and introduces considerable computation, thus entailing degraded accuracy as well as increased complexity of the model. To address this issue, we propose a Memory Optimization Network for Enhanced Vision-Language Target Tracking (VLMOTrack). Specifically, we design a two-stage token filtering mechanism (TTFM). It employs attention hierarchical for adjusting target weights to achieve progressive screening. The first stage uses the stable initial template features for coarse-grained filtering background information. The second stage finely eliminates non-target tokens by utilizing steady language and rich historical visual prompts. Thus, the model obtains precise target feature representations. Furthermore, we construct an adaptive memory prompt generation mechanism (AMPGM). It employs a convolution-based selection strategy to identify reliable frames, which are encoded into a key-value pair structure through a lightweight encoder and convolutional block for storage in a memory bank. Subsequent frames adaptively aggregate values, thereby generating accurate and customized memory prompts. Finally, we propose a memory prompt propagation strategy. The memory prompts are fused with search features through a convolution block to enhance the temporal consistency and discriminability. It then operates as described in the TTFM, thereby propagating historical prompts to subsequent frames for prediction and inference. We conduct experiments on the TNL2K, LaSOT, OTB99-Lang and LaSOText datasets. The experimental results demonstrate that VLMOTrack possesses favorable competitiveness.
Keywords:
Object tracking
Vision-language
Memory network
Prompt learning

Journal

T
The Journal of Supercomputing
IF:
0
Papers:
647
Citations:
0

Organization

E
electrical and information engineering
Scholars:
108
Papers: 39
Citations: 0
C
computer science and technology
Scholars:
435
Papers: 173
Citations: 0
S
software engineering
Scholars:
58
Papers: 37
Citations: 0
S
Shenzhen Institute for Advanced Study
Scholars:
48
Papers: 21
Citations: 0
researcher View more organizations
Cited Papers

Cited Papers

errShare
errSave
Autoregressive Visual Tracking
err2023-06-01
err0
PREAI
errXing Wei; Yifan Bai; Yongchao Zheng; Dahu Shi; Yihong Gong
errShare
errSave
Transformer vision-language tracking via proxy token guided cross-modal fusion
err2023-04-01
err13
PREAI
errZhao, Haojie; Wang, Xiao; Wang, Dong; Lu, Huchuan; Ruan, Xiang
errShare
errSave
Tracking by Natural Language Specification
err2017-07-01
err0
errOAAI
errZhenyang Li; Ran Tao; Efstratios Gavves; Cees G. M. Snoek; Arnold W. M. Smeulders
errShare
errSave
CBAM: Convolutional Block Attention Module
err2018-10-06
err0
PREAI
errSanghyun Woo; Jongchan Park; Joon-Young Lee; In So Kweon
errShare
errSave
Target–distractor memory joint tracking algorithm via Credit Allocation Network
err2024-03-01
err0
PREAI
errZhang,Huanlong; Wang,Panyun; Chen,Zhiwu; Zhang,Jie; Li,Linwei
errShare
errSave
errShare
errSave
researcher View more