Return
Align What Truly Matters: Pedestrian-Relevant Hierarchical Parsing Network for Text-Based Person Retrieval
J
Y
P
朱
H
L
Z
DOI:10.1049/cit2.70163.png)
Abstract
En 中文
Fine-grained alignment is crucial for text-based person retrieval, which searches for relevant pedestrian images using a text query. However, background clutter and semantically vacuous words can cause interference and misalignment, hindering retrieval performance. Existing methods employ masked language modelling to randomly mask and predict words, or leverage multimodal large language models to generate diverse and pure samples. However, these methods rely on external tools and are not sufficiently stable or controllable. In this paper, we propose a pedestrian-relevant hierarchical parsing (PHP) module to extract well-aligned fine-grained visual and textual features for alignment. First, we design a coarse relevant feature mapping (CRFM) module, which uses learnable unified tokens to project both modalities into a shared low-dimensional space, enabling coarse-level semantic filtering. Next, we design an expert-driven feature parsing (EFP) module, which integrates the representational power of a mixture of experts with a modality-aware gating mechanism to uncover deep semantic associations between text and image features. Extensive experiments on three public large-scale datasets demonstrate that our approach outperforms existing state-of-the-art methods, achieving, in particular, a 5.68% improvement in mAP and a 4.81% boost in R1 accuracy on the CUHK-PEDES dataset. Our code is available at https://github.com/Tedysu0916/php_V1.
Keywords:
cross-modal retrieval
mixture of experts
pedestrian feature parsing
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
7.3
Papers:
649
Citations:
2.4K
