arrow
Return

Enhancing visual contextual semantic information for image captioning

delete2025-05-11
delete0
PRE
AI
R
Ronggui Wang
S
Shuo Li
L
Lixia Xue
杨娟 (Juan Yang) *
DOI:10.1007/s13042-025-02634-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning is a fundamental task in the multimodal field, where the goal is to transform images into coherent text through deep network processing. Simple grid features often exhibit subpar performance due to their lack of contextual information and the presence of excessive noise. This paper aims to address these issues by enhancing fine-grained grid features with contextual semantic information within the original transformer model framework. We proposed a novel Dilated Attention Fusion Transformer (DAFT). Firstly, we integrate semantic segmentation features through a Feature Fusion Module Based on Cross-Attention to capture object-related information comprehensively. We then propose a novel multi-scale multi-head sparse attention mechanism based on grids, which improves granularity while reducing unnecessary noise and computational costs. Additionally, we employ a weighted residual connection method to fuse multi-layer information, generating richer representations. Extensive experiments on MS-COCO dataset demonstrate the effectiveness of our DAFT, with improvements of CIDEr from 133.2 to 135.7%, achieving significantly improved performance over the baseline. The source code is available at https://github.com/lishuo19981027/DAFT.
Keywords:
Image captioning
Dilated attention
Transformer

Journal

International Journal of Machine Learning and Cybernetics cover
International Journal of Machine Learning and Cybernetics
IF:
2.7
Papers:
3.1K
Citations:
5.6K

Organization

H
hefei univ technol
Scholars:
1.9K
Papers: 759
Citations: 248