Return
CLDTracker: A Comprehensive Language Description for visual Tracking
DOI:10.1016/j.inffus.2025.103374.png)
Abstract
En 中文
• Introduce comprehensive bag of textual descriptions for VOT tracking. • Provide a comprehensive bag of textual descriptions for six VOT datasets. • Propose TTFUM to update target text features over time. • Fuse visual and textual features using attention-based correlation. • Evaluate CLDTrack on six benchmarks against 38 SOTA trackers.
Keywords:
Multi-modal fusion
Vision–Language Models (VLMs)
Visual Object Tracking (VOT)
Journal
IF:
15.5
Papers:
4.1K
Citations:
2.7W
Organization
No organization information available

