arrow
返回

An efficient automated image caption generation by the encoder decoder model

delete2024-01-22
delete1
PRE
AI
K
Khustar Ansari *
P
Priyanka Srivastava
DOI:10.1007/s11042-024-18150-xdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Image caption generation is becoming one of the hot research topics and attracts various researchers. It is a complex process because it utilizes both NLP (natural language processing) and computer vision approaches for generating the tasks. A range of strategies are available for image captioning that connect the visual material with everyday language, such as explaining images with textual descriptions. Pre-trained classification networks like CNN and RNN-based neural network models are used in the literature to encrypt visual data. Even though various literature works have analyzed outstanding image caption techniques, they still lack in providing better performance for diverse databases. To overcome such issues, this research work presents an automated optimization deep learning model for image caption generation. Initially, the input image is pre-processed, and then the encoder decoder-based structure is utilized for extracting the visual features and caption generation. On the encoder side, the pre-trained ResNet 101 (residual network) is used to extract the visual features, and the SA- Bi-LSTM (self-attention with bi-directional Long Short-Term Memory) is used to generate the caption on the decoder side. In addition, an optimization model CA (Chimp algorithm) is used to improve detection performance in caption generation. The proposed encoder-decoder model is tested on benchmark datasets like Flickr8k, Flickr30k and COCO. Further, this model attained better BLEU and ribes scores of 0.8595 and 0.3531 on the Flickr8k dataset. Thus, the proposed SA-BiLSTM model achieved a significant performance in image caption generation.
Keyword:
Chimp Algorithm
Deep Learning Models
Decoder
Encoder
Image Caption Generation
Visual Features

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
1.9W
被引数:
3.2W

机构

B
birsa institute of technology (bit sindri)
学者数:
120
论文数: 81
被引数: 0
引用论文

引用论文

Image caption generation using Visual Attention Prediction and Contextual Spatial Relation Extraction
err2023-02-08
err12
errOAAI
errSasibhooshan, Reshmi; Kumaraswamy, Suresh; Sasidharan, Santhoshkumar
err分享
err收藏
Rabbit meat in the east of Algeria: motivation and obstacles to consumption
err2020-12-30
err0
errOAAI
errIbtissem Sanah; Samira Becila; Fairouz Djeghim; Abdelghani Boudjellal
err分享
err收藏
err分享
err收藏
Aflatoxinas em produtos à base de milho comercializados no Brasil e riscos para a saúde humana
err2006-06-01
err0
errOAAI
errKassia Ayumi Segawa do Amaral; Gabriel Bassaga Nascimento; Beatriz Leiko Sekiyama; Vanderly Janeiro; Miguel Machinski Jrs
err分享
err收藏
Hydrophobic n-Alkyl-N-isoquinolinium Salts: Ionic Liquids and Low Melting Solids
err2009-07-23
err0
PREAI
errAnn E. Visser; Jonathan G. Huddleston; John D. Holbrey; W. Matthew Reichert; Richard P. Swatloski; Robin D. Rogers
err分享
err收藏
Image captioning model using attention and object features to mimic human image understanding
err2022-02-14
err25
errOAAI
errAl-Malla, Muhammad Abdelhadie; Jafar, Assef; Ghneim, Nada
err分享
err收藏
Neural Image Caption Generation with Weighted Training and Reference
err2018-08-08
err34
errOAAI
errDing, Guiguang; Chen, Minghai; Zhao, Sicheng; Chen, Hui; Han, Jungong; Liu, Qiang
err分享
err收藏
Dual Global Enhanced Transformer for image captioning
err2022-04-01
err63
PREAI
errXian, Tiantao; Li, Zhixin; Zhang, Canlong; Ma, Huifang
err分享
err收藏
err分享
err收藏
学者 查看更多内容