Return
Implicit Alignment with Complementary Information for Text-based Person Re-identification
DOI:10.1016/j.knosys.2026.116029.png)
Abstract
En 中文
Text-based person re-identification (TBReID) aims to identify target pedestrians across different cameras based on given text clues. This task is challenging owing to the inherent cross-modal heterogeneity between images and text, which often results in granularity-imbalanced feature representations. Although existing implicit alignment methods mitigate granularity imbalance and avoid semantic bias by employing learnable units, they overlook the structural semantic asymmetry between modalities, such as differences in body proportion and posture. To address these limitations, we propose Implicit Alignment framework with Complementary Information (IACI), a novel framework that enhances textual representations using structural semantics to achieve better cross-modal alignment. The IACI framework comprises three key modules: a Feature Compensation Fine-grained Alignment module, which leverages sketch generation to enrich textual structural semantics and enforces fine-grained alignment constraints; a Relationship-based Noise Filtering module, which employs contrastive noise embedding to reduce interference from image background noise and redundant textual descriptions; and a Prototype-based Granularity Balancing module, which utilizes shared dynamic prototypes to align modalities using temperature-scaled attention. Extensive experiments on three mainstream datasets demonstrate that our method achieves state-of-the-art performance in TBReID. The underlying code will be released upon acceptance.
Keywords:
Text-based person re-identification
cross-modal alignment
structural semantics
granularity balancing
noise filtering
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

