arrow
Return

An efficient automated image caption generation by the encoder decoder model

delete2024-01-22
delete1
PRE
AI
K
Khustar Ansari *
P
Priyanka Srivastava
DOI:10.1007/s11042-024-18150-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image caption generation is becoming one of the hot research topics and attracts various researchers. It is a complex process because it utilizes both NLP (natural language processing) and computer vision approaches for generating the tasks. A range of strategies are available for image captioning that connect the visual material with everyday language, such as explaining images with textual descriptions. Pre-trained classification networks like CNN and RNN-based neural network models are used in the literature to encrypt visual data. Even though various literature works have analyzed outstanding image caption techniques, they still lack in providing better performance for diverse databases. To overcome such issues, this research work presents an automated optimization deep learning model for image caption generation. Initially, the input image is pre-processed, and then the encoder decoder-based structure is utilized for extracting the visual features and caption generation. On the encoder side, the pre-trained ResNet 101 (residual network) is used to extract the visual features, and the SA- Bi-LSTM (self-attention with bi-directional Long Short-Term Memory) is used to generate the caption on the decoder side. In addition, an optimization model CA (Chimp algorithm) is used to improve detection performance in caption generation. The proposed encoder-decoder model is tested on benchmark datasets like Flickr8k, Flickr30k and COCO. Further, this model attained better BLEU and ribes scores of 0.8595 and 0.3531 on the Flickr8k dataset. Thus, the proposed SA-BiLSTM model achieved a significant performance in image caption generation.
Keywords:
Chimp Algorithm
Deep Learning Models
Decoder
Encoder
Image Caption Generation
Visual Features

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
2.0W
Citations:
3.2W

Organization

B
birsa institute of technology (bit sindri)
Scholars:
120
Papers: 81
Citations: 0
Cited Papers

Cited Papers

Image caption generation using Visual Attention Prediction and Contextual Spatial Relation Extraction
err2023-02-08
err12
errOAAI
errSasibhooshan, Reshmi; Kumaraswamy, Suresh; Sasidharan, Santhoshkumar
errShare
errSave
Rabbit meat in the east of Algeria: motivation and obstacles to consumption
err2020-12-30
err0
errOAAI
errIbtissem Sanah; Samira Becila; Fairouz Djeghim; Abdelghani Boudjellal
errShare
errSave
Deep Learning Approaches Based on Transformer Architectures for Image Captioning Tasks
err2022-01-01
err21
errOAAI
errCastro, Roberto; Pineda, Israel; Lim, Wansu; Morocho-Cayamcela, Manuel Eugenio
errShare
errSave
Aflatoxinas em produtos à base de milho comercializados no Brasil e riscos para a saúde humana
err2006-06-01
err0
errOAAI
errKassia Ayumi Segawa do Amaral; Gabriel Bassaga Nascimento; Beatriz Leiko Sekiyama; Vanderly Janeiro; Miguel Machinski Jrs
errShare
errSave
Hydrophobic n-Alkyl-N-isoquinolinium Salts: Ionic Liquids and Low Melting Solids
err2009-07-23
err0
PREAI
errAnn E. Visser; Jonathan G. Huddleston; John D. Holbrey; W. Matthew Reichert; Richard P. Swatloski; Robin D. Rogers
errShare
errSave
Image captioning model using attention and object features to mimic human image understanding
err2022-02-14
err25
errOAAI
errAl-Malla, Muhammad Abdelhadie; Jafar, Assef; Ghneim, Nada
errShare
errSave
Neural Image Caption Generation with Weighted Training and Reference
err2018-08-08
err34
errOAAI
errDing, Guiguang; Chen, Minghai; Zhao, Sicheng; Chen, Hui; Han, Jungong; Liu, Qiang
errShare
errSave
Dual Global Enhanced Transformer for image captioning
err2022-04-01
err63
PREAI
errXian, Tiantao; Li, Zhixin; Zhang, Canlong; Ma, Huifang
errShare
errSave
errShare
errSave
researcher View more