Return
Implicit visual knowledge enhanced zero-shot image captioning
DOI:10.1016/j.eswa.2025.130862.png)
Abstract
En 中文
• We introduce implicit visual knowledge to bridge the gap between CLIP and language models. • A Visual Knowledge Extraction (VKE) module extracts image-related knowledge using GPT-2. • Two integration methods enhance caption accuracy and diversity. • Our framework achieves a 15.9 % improvement in caption diversity on MSCOCO.
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W
Organization
No organization information available

