arrow
Return

Graph-based image captioning with semantic and spatial features

delete2025-04-01
delete0
PRE
AI
M
Mohammad Javad Parseh *
S
Saeed Ghadiri
DOI:10.1016/j.image.2025.117273delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning is a challenging task of image processing that aims to generate descriptive and accurate textual descriptions for images. In this paper, we propose a novel image captioning framework that leverages the power of spatial and semantic relationships between objects in an image, in addition to traditional visual features. Our approach integrates a pre-trained model, RelTR, as a backbone for extracting object bounding boxes and subjectpredicate-object relationship pairs. We use these extracted relationships to construct spatial and semantic graphs, which are processed through separate Graph Convolutional Networks (GCNs) to obtain high-level contextualized features. At the same time, a CNN model is employed to extract visual features from the input image. To merge the feature vectors seamlessly, our approach involves using a multi-modal attention mechanism that is applied separately to the feature maps of the image, the nodes of the semantic graph, and the nodes of the spatial graph during each time step of the LSTM-based decoder. The model concatenates the attended features with the word embedding at the respective time step and fed into the LSTM cell. Our experiments demonstrate the effectiveness of our proposed approach, which competes closely with existing state-of-the-art image captioning techniques by capturing richer contextual information and generating accurate and semantically meaningful captions. (c) 2025 Elsevier Inc. All rights reserved.
Keywords:
Image captioning
Semantic Graph
Spatial graph
Attention mechanism

Journal

S
Signal Processing and Image Communication
IF:
2.7
Papers:
2.8K
Citations:
4.2K

Organization

No organization information available