1
Return

Dual Graph Network Hashing for Cross-Modal Retrieval

delete2026-07-06
delete0
PRE
AI
S
Shuang Zhang
Y
Yue Wu
L
Lei Shi
F
Feifei Kou
H
Huilong Jin
P
Pengfei Zhang
W
Weiping Ding
M
Mingying Xu
M
Muhammet Deveci
DOI:10.1109/tkde.2026.3710651delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Hashing algorithms represent data by generating compact binary hash codes, enabling efficient cross-modal similarity search and significantly improving the storage efficiency and retrieval performance of image-text data. However, because traditional hashing methods typically separate image-text feature extraction from hash learning, the feature extraction module struggles to adaptively update based on training feedback, further limiting the performance of cross-modal retrieval in real-world scenarios. To address this issue, deep learning has been introduced to cross-modal hashing, enabling end-to-end joint optimization and tightly integrating feature extraction and hash learning, significantly improving retrieval performance. However, because existing deep learning methods often use a fixed weight distribution when processing samples, they ignore the modal differences between text and visual features when fusing them. This leads to an inadequate fused representation and difficulty achieving optimal modality alignment during hash code generation. To address this issue, we propose a Dual Graph Network Hashing (DGNH) algorithm that dynamically adjusts the weight distribution between visual and text features through an adaptive attention mechanism, ensuring better modality fusion during hash code generation. Specifically, we design a novel framework that combines a graph convolutional neural network (GCN) with a graph attention network (GAT) to construct a label classifier for generating labels and enhancing cross-modal feature representation. This approach improves feature discrimination by capturing the hierarchical relationships and co-occurrence patterns of labels through a carefully constructed label association graph. Furthermore, we introduced a pre-trained model combining CLIP and the Transformer to further enhance the overall feature representation. During the optimization phase, we employed a contrastive triplet loss function coupled with novel regularization constraints for quantization and optimization, thus effectively reducing information loss during discretization and ensuring the generated hash codes are more compact and efficient. Experimental results on three public datasets, MS-COCO, NUS-WIDE, and MIRFlickr-25 K, demonstrate that the proposed method outperforms existing methods in both retrieval accuracy and efficiency, thereby validating its effectiveness and superiority.
Keywords:
Cross-modal hashing retrieval
adaptive attention mechanisms
graph convolutional network
graph attention network

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

N
north china university of technology
Scholars:
779
Papers: 340
Citations: 0
A
anhui university of science and technology
Scholars:
1.2K
Papers: 447
Citations: 0
N
nantong university
Scholars:
3.4K
Papers: 1.0K
Citations: 0
C
communication university of china
Scholars:
404
Papers: 211
Citations: 0
National Defence University cover
National Defence University
Scholars:
39
Papers: 51
Citations: 22
B
H
Hebei Normal University
Scholars:
6.2K
Papers: 3.4K
Citations: 9
Cited Papers

Cited Papers

Citing Papers

Citing Papers