返回
Cross-Modal Generation and Pair Correlation Alignment Hashing
DOI:10.1109/TITS.2022.3221787.png)
摘要
En 中文
Cross-modal hashing is an effective cross-modal retrieval approach because of its low storage and high efficiency. However, most existing methods mainly utilize pre-trained networks to extract modality-specific features, while ignore the position information and lack information interaction between different modalities. To address those problems, in this paper, we propose a novel approach, named cross-modal generation and pair correlation alignment hashing (CMGCAH), which introduces transformer to exploit position information and utilizes cross-modal generative adversarial networks (GAN) to boost cross-modal information interaction. Concretely, a cross-modal interaction network based on conditional generative adversarial network and pair correlation alignment networks are proposed to generate cross-modal common representations. On the other hand, a transformer-based feature extraction network (TFEN) is designed to exploit position information, which can be propagated to text modality and enforce the common representation to be semantically consistent. Experiments are performed on widely used datasets with text-image modalities, and results show that the proposed method achieved competitive performance compared with many existing methods.
Keyword:
Cross-modal hashing
cross-modal generation
correlation alignment
position semantic information
cross-modal interaction
期刊
IF:
8.4
论文数:
9.7K
被引数:
6.3W
机构
引用论文
Preischemic Administration of Flunarizine or Phencyclidine Reduces Local Cerebral Glucose Utilization in Rat Hippocampus Seven Days after Ischemia
Pharmacology
IF0
Clustering-driven Deep Adversarial Hashing for scalable unsupervised cross-modal retrieval用于可扩展的无监督跨模态检索的聚类驱动的深度对抗哈希
NEUROCOMPUTING
IF6.5

