arrow
返回

Enhance Composed Image Retrieval via Multi-Level Collaborative Localization and Semantic Activeness Perception

delete2024-01-01
delete0
PRE
AI
G
Gangjian Zhang
韦世奎 (Shikui Wei) *
H
Huaxin Pang
S
Shuang Qiu
赵耀 (Yao Zhao)
DOI:10.1109/TMM.2023.3273466delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Composed image retrieval (CIR) is an emerging and challenging research task that combines two modalities, a reference image, and a modification text, into one query to retrieve the target image. In online shopping scenarios, the user would use the modification text as feedback to describe the difference between the reference and the desired image. In order to handle the task, there must be two main problems needed to be addressed. One is the localization problem: how to precisely find those spatial areas of the image mentioned by the text. The other is the modification problem: how to effectively modify the image semantics based on the text. However, existing methods merely fuse information coarsely from the two-modality, while the accurate spatial and semantic correspondence between these two heterogeneous features tends to be neglected. Therefore, image details cannot be precisely located and modified. To this end, we consider integrating information from the two modalities more accurately from spatial and semantic aspects. Thus, we propose an end-to-end framework for the CIR task, which contains three key components, i.e., Multi-level Collaborative Localization module (MCL), Differential Semantics Discrimination module (DSD), and Image Difference Enhancement constraints (IDE). Specifically, to solve the localization problem, MCL precisely locates the text to the image areas by collaboratively using text positioning information on multiple image layers. For the modification problem, DSD builds a distribution to evaluate the modification possibility of each image semantic dimension, and IDE effectively learns the modification patterns of text against image embedding based on the distribution. Extensive experiments on three datasets show that the proposed method achieves outstanding performance against the SOTA methods.
Keyword:
Semantics
Location awareness
Task analysis
Image retrieval
Training
Collaboration
Transformers
Composed image retrieval
multi-modal fusion and embedding
multi-modal representation learning
multi-modal retrieval
image retrieval

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

B
Beijing Jiaotong University
学者数:
2.2W
论文数: 1.7W
被引数: 1.2W
引用论文

引用论文

Direct replication of Gervais & Norenzayan (2012): No evidence that analytic thinking decreases religious belief
err2017-02-24
err0
errOAAI
errClinton Sanchez; Brian Sundermeier; Kenneth Gray; Robert J. Calin-Jageman
err分享
err收藏
Expression in baculovirus vector system of the nucleocapsid protein gene of rinderpest virus
err1993-07-01
err0
PREAI
errH. Kamata; S. Ohkubo; M. Sugiyama; Y. Matsuura; Y. Kamata; K. Tsukiyama-Kohara; K. Imaoka; C. Kai; Y. Yoshikawa; K. Yamanouchi
err分享
err收藏
WT1 Promotes Invasion of NSCLC via Suppression of CDH1
err2013-09-01
err0
errOAAI
errChen Wu; Weiyou Zhu; Jing Qian; Shaohua He; Changping Wu; Yijiang Chen; Yongqian Shu
err分享
err收藏
Transient Convective Effects on Diffusion Measurements in Liquids
err2003-04-01
err0
PREAI
errJohn P. Kizito; J. Iwan D. Alexander; R. Michael Banish
err分享
err收藏
The Application of Cyclobutane Derivatives in Organic Synthesis
err2003-03-15
err0
PREAI
errJan C. Namyslo; Dieter E. Kaufmann
err分享
err收藏
学者 查看更多内容