arrow
返回

ETransCap: efficient transformer for image captioning

delete2024-08-27
delete0
PRE
AI
A
Albert Mundu *
S
Satish Kumar Singh
S
Shiv Ram Dubey
DOI:10.1007/s10489-024-05739-wdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Image captioning is a challenging task in computer vision that automatically generates a textual description of an image by integrating visual and linguistic information, as the generated captions must accurately describe the image's content while also adhering to the conventions of natural language. We adopt the encoder-decoder framework employed by various CNN-RNN-based models for image captioning in the past few years. Recently, we observed that the CNN-Transformer-based models have achieved great success and surpassed traditional CNN-RNN-based models in the area. Many researchers have concentrated on Transformers, exploring and uncovering its vast possibilities. Unlike conventional CNN-RNN-based models in image captioning, transformer-based models have achieved notable success and offer the benefit of handling longer input sequences more efficiently. However, they are resource-intensive to train and deploy, particularly for large-scale tasks or for tasks that require real-time processing. In this work, we introduce a lightweight and efficient transformer-based model called the Efficient Transformer Captioner (ETransCap), which consumes fewer computation resources to generate captions. Our model operates in linear complexity and has been trained and tested on MS-COCO dataset. Comparisons with existing state-of-the-art models show that ETransCap achieves promising results. Our results support the potential of ETransCap as a good approach for image captioning tasks in real-time applications. Code for this project will be available at https://github.com/albertmundu/etranscap.
Keyword:
Deep learning
Natural language processing
Image captioning
Scene understanding
Transformers
Efficient transformers

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

I
Indian Institute of Information Technology Allahabad
学者数:
857
论文数: 627
被引数: 833
引用论文

引用论文

err分享
err收藏
Depth-of-Field of the Accommodating Eye
err2014-10-01
err0
errOAAI
errPaula Bernal-Molina; Robert Montés-Micó; Richard Legras; Norberto López-Gil
err分享
err收藏
err分享
err收藏
Remote sensing of fish-processing in the Sundarbans Reserve Forest, Bangladesh: an insight into the modern slavery-environment nexus in the coastal fringe
err2020-09-17
err0
errOAAI
errBethany Jackson; Doreen S. Boyd; Christopher D. Ives; Jessica L. Decker Sparks; Giles M. Foody; Stuart Marsh; Kevin Bales
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容