arrow
Return

Image-text matching using multi-subspace joint representation

delete2023-01-07
delete2
PRE
AI
孙昊 (Hao Sun)
秦小麟 (Xiaolin Qin) *
刘小靖 (Xiaojing Liu)
DOI:10.1007/s00530-022-01038-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Joint representation learning has been an attractive way to solve image-text retrieval problem due to its efficiency on both time and storage. On the one hand, the most classical methods model the joint semantic subspace with respect to only semantic relationship at the level of holistic images and sentences, which fails to explore the fine-grained semantic relationship. On the other hand, visual-linguistic pretrain-finetune scheme has achieved an impressive performance in many downstream image-text tasks, but more storage cost and computation burden still notably limit their use in some real applications which have strict requirement for low storage cost or rapid response to image acquired in real-time, such as retrieving textual description of some just taken photos on a storage-limited mobile platform. To mitigate the problem above, we proposed a lightweight cross-modal retrieval for learning the joint representation. Instead of only modeling whole joint semantic space, the proposed model captures semantic relationship in multiple subspaces. Specially, we treat the retrieval problem as not only a ranking process but also a decision process, and propose an entropy-based constraint to preserve as much hierarchy-aware information as possible across various semantic subspaces. Experiments show that the proposed method can achieve competitive performance when compared with the state-of-the-art joint representation learning methods on two publicly available datasets.
Keywords:
Joint representation
Image-text retrieval
Multi-subspace learning
Cross-modal matching

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

No organization information available