arrow
Return

Dynamic sparse and weight allocation-based text-driven person retrieval

delete2025-09-24
delete0
PRE
AI
S
Shuren Zhou *
Q
Qihang Zhou
J
Jiao Liu
DOI:10.1016/j.imavis.2025.105737delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose a progressive enhancement for a multimodal integration strategy that ensures that information at different levels of detail can be effectively integrated and aligned, thus improving the model’s ability to understand complex data. • We adopt a global two-way match filtering method, which can effectively mitigate the interference of false matching pairs of image text, so as to select the image text with a high matching degree for training. • We introduce a fine-grained dynamic sparse mask modeling approach, which not only extracts the key elements in the features but also mines the image text for fine-grained relationships, thus improving the retrieval performance.

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.1K
Citations:
6.7K

Organization

C
changsha university of science and technology
Scholars:
3.3K
Papers: 1.2K
Citations: 0
U
university of south china
Scholars:
1.4W
Papers: 6.8K
Citations: 8
Cited Papers

Cited Papers

No cited papers available