Return
Multiple Local-Guided Modality-Shared Feature Learning Transformer for Visible-Infrared Person Re-Identification
X
H
S
DOI:10.1109/JIOT.2026.3673575.png)
Abstract
En 中文
Visible–infraredperson re-identification (VI-ReID) aims to identify the same person in heterogeneous images captured by different spectral cameras. Due to the different imaging mechanisms of heterogeneous cameras, there are significant modality differences between cross-modal images. Directly extracting modality-shared features from heterogeneous images can make it difficult to align local discriminative information between different modalities, leading to local semantic inconsistency. To address this, we propose a novel multiple local-guided modality-shared feature learning transformer (MLMT) model to mine modality-shared local semantic similarity information, thereby facilitating the effective alignment of cross-modal local semantics. First, we design a parameter-free random modality gap bridging (RMGB) module, which uses a low-complexity strategy involving random local region grayscale transformation and inverse grayscale transformation. This approach reduces the model’s sensitivity to modality changes from a cross-modal local perspective, thus smoothly and efficiently bridging the inherent differences between the two modalities. Second, to better learn the shared local discriminative information, we design a multiple local-guided feature learning (ML) strategy. This strategy leverages learnable global and local tokens and uses orthogonal constraint learning to focus on distinct discriminative information in overlapping local regions, thereby exploring shared local similarities between different modalities. Furthermore, we introduce a similarity inference reinforcement (SIR) module, which uses high similarity information between person images to optimize the distance matrix, thereby improving matching performance. Extensive experiments on two benchmark datasets for the VI-ReID task demonstrate the effectiveness of the MLMT method.
Keywords:
Cross-modality
feature alignment
Internet of Things (IoT)
modality-shared
person re-identification (Re-ID)
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

