arrow
Return

Multi-modal visual tracking based on textual generation

delete2024-12-01
delete0
PRE
AI
J
Jiahao Wang
刘芳 (Fang Liu) *
L
Licheng Jiao
H
Hao Wang
S
Shuo Li
L
Lingling Li
陈璞花 (Puhua Chen)
X
Xu Liu
DOI:10.1016/j.inffus.2024.102531delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multi-modal tracking has garnered significant attention due to its wide range of potential applications. Existing multi-modal tracking approaches typically merge data from different visual modalities on top of RGB tracking. However, focusing solely on the visual modality is insufficient due to the scarcity of tracking data. Inspired by the recent success of large models, this paper introduces a Multi-modal Visual Tracking Based on Textual Generation (MVTTG) approach to address the limitations of visual tracking, which lacks language information and overlooks semantic relationships between the target and the search area. To achieve this, we leverage large models to generate image descriptions, using these descriptions to provide complementary information about the target's appearance and movement. Furthermore, to enhance the consistency between visual and language modalities, we employ prompt learning and design a Visual-Language Interaction Prompt Manager (V-L PM) to facilitate collaborative learning between visual and language domains. Experiments conducted with MVTTG on multiple benchmark datasets confirm the effectiveness and potential of incorporating image descriptions in multi-modal visual tracking.
Keywords:
Multi-modal tracking
Image descriptions
Visual and language modalities
Prompt learning

Journal

Information Fusion cover
Information Fusion
IF:
15.5
Papers:
4.1K
Citations:
2.7W

Organization

X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K