arrow
Return

Mamba-Caption: Long-Range Sequence Modelling for Efficient and Accurate Image Captioning

delete2025-10-24
delete0
delete
OA
AI
T
Tariq Shahzad
M
Muhammad Aoun *
T
Tehseen Mazhar *
M
Muhammad Usman Tariq
K
Khmaies Ouahada
H
Habib Hamam
DOI:10.1016/j.array.2025.100538delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Image captioning has been a problem in vision–language research for a long time. Long-range dependencies and efficiency are challenges for the standard models, such as recurrent neural networks (RNNs) and Transformers. To overcome this, we present Mamba-Caption, an efficient sequence processing model that replaces attention mechanisms with selective state-space modelling. The core novelty is a Mamba-based decoder that substitutes self-attention with selective state-space updates, enabling linear-time caption generation while preserving long-range token dependencies; this decoder is a drop-in language-side component that conditions on a convolutional neural network (CNN) image embedding without domain-specific heuristics. Our model utilizes a CNN encoder, a token embedding layer, and a Mamba-based decoder; the decoder is trained using teacher forcing with a cross-entropy objective. Our model outperforms baselines on all standard metrics when evaluated on the Flickr30k dataset, achieving a Bilingual Evaluation Understudy (BLEU-1) score of 0.83, a Metric for Evaluation of Translation with Explicit ORdering (METEOR) score of 0.79, a Recall-Oriented Understudy for Gisting Evaluation—Longest Common Subsequence (ROUGE-L) score of 0.73, and a Consensus-based Image Description Evaluation (CIDEr) score of 1.30. We further contextualize efficiency via a qualitative/complexity discussion and ablation framing that isolates decoder-side design choices, reinforcing that the gains in efficiency do not sacrifice accuracy. Mamba-Caption can be applied to real-world captioning tasks due to its high efficiency and generalizability.
Keywords:
CNN
Mamba
CIDEr
Meteor
Blue Score
RNN
Res Net
Encoder
Decoder
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Array cover
Array
IF:
4.5
Papers:
701
Citations:
1.2K

Organization

A
Abu Dhabi University
Scholars:
1.2K
Papers: 1.3K
Citations: 1.7K
G
Ghazi University
Scholars:
47
Papers: 32
Citations: 711
U
University of Johannesburg
Scholars:
6.8K
Papers: 6.8K
Citations: 1.2W
researcher View more organizations