返回
Image-text bidirectional learning network based cross-modal retrieval
DOI:10.1016/j.neucom.2022.02.007.png)
摘要
En 中文
The problem of cross-modal retrieval has attracted significant attention in the cross-media retrieval community. One key challenge of cross-modal retrieval is to eliminate the heterogeneous gap between different patterns. The existing numerous cross-modal retrieval approaches tend to jointly construct a common subspace, while these methods fail to consider mutual influence between modalities sufficiently during the whole training process. In this paper, we propose a novel image-text Bidirectional Learning Network (BLN) based cross-modal retrieval method. The method constructs a common representation space and directly measures the similarity of heterogeneous data. More specifically, a multi-layer supervision network is proposed to learn the cross-modal relevance of the generated representations. Moreover, a bidirectional crisscross loss function is proposed to preserve the modal invariance with the bidirectional learning strategy in the common representation space. The loss functions of discriminant consistency and the bidirectional crisscross loss are integrated into an objective function which aims to minimize the intra-class distance and maximize the inter-class distance. Comprehensive experimental results on four widely-used databases show that the proposed method is effective and superior to the existing cross-modal retrieval methods. (c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Cross-modal retrieval
bidirectional learning network
common representation space
discriminant consistency loss
bidirectional crisscross loss
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
CM-GANs: Cross-modal Generative Adversarial Networks for Common Representation Learning二1212-GANs: 用于通用表示学习的跨模态生成对抗网络
The effect of decomposition of beta-phase Zr-20 at% Nb on hydrogen partitioning with alpha-zirconium

