Return
Text-based person search via fine-grained cross-modal semantic alignment
F
J
Y
X
DOI:10.1016/j.image.2026.117478.png)
Abstract
En 中文
Existing text-based person search methods face challenges in handling complex cross-modal interactions, often failing to capture subtle semantic nuances. To address this, we propose a novel Fine-grained Cross-modal Semantic Alignment (FCSA) framework that enhances accuracy and robustness in text-based person search. FCSA introduces two key components: the Cross-Modal Reconstruction Strategy (CMRS) and the Saliency-Guided Masking Mechanism (SGMM). CMRS facilitates feature alignment by leveraging incomplete visual and textual features, promoting bidirectional reasoning across modalities, and enhancing fine-grained semantic understanding. SGMM further refines performance by dynamically focusing on salient visual patches and critical text tokens, thereby improving discriminative region perception and image-text matching precision. Our approach outperforms existing state-of-the-art methods, achieving mean Average Precision (mAP) scores of 69.72%, 43.78% and 48.78% on CUHK-PEDES, ICFG-PEDES, and RSTPReid, respectively. Source code is at https://github.com/flychen321/FCSA.
Keywords:
Text-based person search
Cross-modal semantic alignment
Saliency-guided masking mechanism
Visual and textual feature
Journal
S
IF:
2.7
Papers:
18
Citations:
0
