arrow
Return

Multiple Local-Guided Modality-Shared Feature Learning Transformer for Visible-Infrared Person Re-Identification

delete2026-03-12
delete0
PRE
AI
X
Xiyuan Wang
H
Huadong Zhang
S
Shuli Cheng
DOI:10.1109/JIOT.2026.3673575delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visible–infraredperson re-identification (VI-ReID) aims to identify the same person in heterogeneous images captured by different spectral cameras. Due to the different imaging mechanisms of heterogeneous cameras, there are significant modality differences between cross-modal images. Directly extracting modality-shared features from heterogeneous images can make it difficult to align local discriminative information between different modalities, leading to local semantic inconsistency. To address this, we propose a novel multiple local-guided modality-shared feature learning transformer (MLMT) model to mine modality-shared local semantic similarity information, thereby facilitating the effective alignment of cross-modal local semantics. First, we design a parameter-free random modality gap bridging (RMGB) module, which uses a low-complexity strategy involving random local region grayscale transformation and inverse grayscale transformation. This approach reduces the model’s sensitivity to modality changes from a cross-modal local perspective, thus smoothly and efficiently bridging the inherent differences between the two modalities. Second, to better learn the shared local discriminative information, we design a multiple local-guided feature learning (ML) strategy. This strategy leverages learnable global and local tokens and uses orthogonal constraint learning to focus on distinct discriminative information in overlapping local regions, thereby exploring shared local similarities between different modalities. Furthermore, we introduce a similarity inference reinforcement (SIR) module, which uses high similarity information between person images to optimize the distance matrix, thereby improving matching performance. Extensive experiments on two benchmark datasets for the VI-ReID task demonstrate the effectiveness of the MLMT method.
Keywords:
Cross-modality
feature alignment
Internet of Things (IoT)
modality-shared
person re-identification (Re-ID)

Journal

IEEE Internet of Things Journal cover
IEEE Internet of Things Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

Organization

C
China Telecom Corporation Ltd.
Scholars:
10
Papers: 10
Citations: 0
X
Xinjiang University of Finance and Economics
Scholars:
260
Papers: 204
Citations: 187
X
xinjiang university
Scholars:
2.9K
Papers: 872
Citations: 0
researcher View more organizations