Return
CLIP-based knowledge projector for image–text matching
DOI:10.1016/j.ipm.2025.104357.png)
Abstract
En 中文
• We put forward a knowledge projector network which regards prior knowledge in CLIP (Radford et al., 2021) as a teacher to guide slot attention generation process. • An adaptive weighted fusion module is used to incorporate global features into slot representations. • An effective similarity calculation method is proposed to compare with fine-grained image–text matching methods. The results indicate that our method outperforms CLIP and the most recent image–text alignment algorithms.
Keywords:
Image–text matching
Multimedia analysis
Slot attention
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W

