arrow
Return

Retrieval-based objects and relations prompt for image captioning

delete2026-03-20
delete0
PRE
AI
J
Jinjing Gu
T
Tianbao Qin *
普园媛 cover
普园媛 (Yuanyuan Pu) *
Z
Zhengpeng Zhao
DOI:10.1016/j.engappai.2026.114518delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning aims to generate natural language descriptions for input images in an open-form manner. To accurately generate descriptions related to the image, a critical step in image captioning is to identify objects and understand their relations within the image. Modern approaches typically capitalize on object detectors or combine detectors with Graph Convolutional Networks (GCNs). However, these models suffer from redundant detection information, difficulty in GCNs construction, and high training costs. To address these issues, the Retrieval-based Objects and Relations Prompt for Image Captioning (RORPCap) is proposed, inspired by the fact that image-text retrieval can provide rich semantic information for input images. RORPCap employs an Objects and Relations Extraction Model to extract object and relation words from the image. These words are then incorporated into predefined prompt templates and encoded as prompt embeddings. Next, a Mamba-based mapping network is designed to quickly map image embeddings extracted by the Contrastive Language-Image Pre-training model into visual-text embeddings. Finally, the resulting prompt embeddings and visual-text embeddings are concatenated to form textual-enriched feature embeddings, which are fed into a GPT-2 model for caption generation. Extensive experiments conducted on the widely used MS-COCO dataset show that the RORPCap requires only 2.6 h under cross-entropy loss training, achieving 120.5% CIDEr score and 22.0% SPICE score on the ‘Karpathy’ test split. RORPCap achieves comparable performance metrics to detector-based and GCN-based models with the shortest training time and demonstrates its potential as an alternative for image captioning. The source code is available at: https://github.com/jinjinggu00/RORPCap .
Keywords:
Image captioning
Object and relation extraction
Prompt-based learning
Mamba-based mapping
Contrastive Language-Image Pre-training

Journal

Engineering Applications of Artificial Intelligence cover
Engineering Applications of Artificial Intelligence
IF:
8
Papers:
5.3K
Citations:
3.5W

Organization

Y
yunnan university
Scholars:
3.8K
Papers: 1.2K
Citations: 0