arrow
返回

Multimodal Composition Example Mining for Composed Query Image Retrieval

delete2024-01-01
delete0
PRE
AI
G
Gangjian Zhang
S
Shikun Li
韦世奎 (Shikui Wei) *
葛仕明 (Shiming Ge)
N
Na Cai
赵耀 (Yao Zhao)
DOI:10.1109/TIP.2024.3359062delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Composed query image retrieval task aims to retrieve the target image in the database by a query that composes two different modalities: a reference image and a sentence declaring that some details of the reference image need to be modified and replaced by new elements. Tackling this task needs to learn a multimodal embedding space, which can make semantically similar targets and queries close but dissimilar targets and queries as far away as possible. Most of the existing methods start from the perspective of model structure and design some clever interactive modules to promote the better fusion and embedding of different modalities. However, their learning objectives use conventional query-level examples as negatives while neglecting the composed query's multimodal characteristics, leading to the inadequate utilization of the training data and suboptimal construction of metric space. To this end, in this paper, we propose to improve the learning objective by constructing and mining hard negative examples from the perspective of multimodal fusion. Specifically, we compose the reference image and its logically unpaired sentences rather than paired ones to create component-level negative examples to better use data and enhance the optimization of metric space. In addition, we further propose a new sentence augmentation method to generate more indistinguishable multimodal negative examples from the element level and help the model learn a better metric space. Massive comparison experiments on four real-world datasets confirm the effectiveness of the proposed method.
Keyword:
Composed query image retrieval
multimodal fusion
multimodal metric learning
hard example mining

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
B
Beijing Jiaotong University
学者数:
2.2W
论文数: 1.7W
被引数: 1.2W
I
institute of information engineering, cas
学者数:
474
论文数: 466
被引数: 0
学者 查看更多机构