arrow
Return

CLIP-based knowledge projector for image–text matching

delete2025-08-23
delete0
PRE
AI
X
Xinfeng Dong
D
Dingwen Zhang
L
Longfei Han
M
Ming Jin *
刘丽 (Li Liu)
J
Junwei Han
DOI:10.1016/j.ipm.2025.104357delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We put forward a knowledge projector network which regards prior knowledge in CLIP (Radford et al., 2021) as a teacher to guide slot attention generation process. • An adaptive weighted fusion module is used to incorporate global features into slot representations. • An effective similarity calculation method is proposed to compare with fine-grained image–text matching methods. The results indicate that our method outperforms CLIP and the most recent image–text alignment algorithms.
Keywords:
Image–text matching
Multimedia analysis
Slot attention

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

B
Beijing Technology and Business University
Scholars:
3.7K
Papers: 1.6K
Citations: 1.6W
N
Northwestern Polytechnical University
Scholars:
4.6W
Papers: 3.7W
Citations: 5.3W
S
shandong normal university
Scholars:
1.0W
Papers: 8.2K
Citations: 3
researcher View more organizations