1
Return

Text-based person search via fine-grained cross-modal semantic alignment

delete2026-01-01
delete1
PRE
AI
F
Feng Chen
J
Jielong He
Y
Yang Liu
X
Xiwen Qu *
DOI:10.1016/j.image.2026.117478delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Existing text-based person search methods face challenges in handling complex cross-modal interactions, often failing to capture subtle semantic nuances. To address this, we propose a novel Fine-grained Cross-modal Semantic Alignment (FCSA) framework that enhances accuracy and robustness in text-based person search. FCSA introduces two key components: the Cross-Modal Reconstruction Strategy (CMRS) and the Saliency-Guided Masking Mechanism (SGMM). CMRS facilitates feature alignment by leveraging incomplete visual and textual features, promoting bidirectional reasoning across modalities, and enhancing fine-grained semantic understanding. SGMM further refines performance by dynamically focusing on salient visual patches and critical text tokens, thereby improving discriminative region perception and image-text matching precision. Our approach outperforms existing state-of-the-art methods, achieving mean Average Precision (mAP) scores of 69.72%, 43.78% and 48.78% on CUHK-PEDES, ICFG-PEDES, and RSTPReid, respectively. Source code is at https://github.com/flychen321/FCSA.
Keywords:
Text-based person search
Cross-modal semantic alignment
Saliency-guided masking mechanism
Visual and textual feature

Journal

S
SIGNAL PROCESSING-IMAGE COMMUNICATION
IF:
2.7
Papers:
18
Citations:
0

Organization

A
anhui university of technology
Scholars:
9.1K
Papers: 5.4K
Citations: 9
Cited Papers

Cited Papers

Citing Papers

Citing Papers