Return
A vision explainability method for image captioning using transformer decoder attention maps
DOI:10.1016/j.mex.2025.103744.png)
Abstract
En 中文
Image Captioning is a crucial task that enables systems to generate descriptive sentences for visual content. Though image captioning systems bloom at the intersection of Computer Vision and Natural Language Processing, these models act mostly as black boxes offering little or no insight into how captions are derived. We present a novel explainable image captioning framework that integrates a Convolutional Neural Network encoder with a Transformer decoder. Attention-based heatmaps are used to explain the visuals offering transparency in the decision making process. The method evaluates captioning quality and interpretability on the MS COCO dataset using BLEU, METEOR, CIDER and SPICE. The method enhances the trustworthiness and transparency, making it reliable for applications like healthcare, education, security, surveillance and forecasting. A reproducible method for integrating visual explainability into image captioning exploring transformer decoder attention maps. The method contributes to the growing body of eXplainable AI (XAI) by addressing the transparency gap in vision-language models Balance performance with interpretability paving the way for more transparent and trustworthy AI systems.
Keywords:
Image captioning
Transformer models
Visual attention maps
Convolutional neural network
Explainable AI
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
1.9
Papers:
290
Citations:
5.6K

