返回
An efficient automated image caption generation by the encoder decoder model
DOI:10.1007/s11042-024-18150-x.png)
摘要
En 中文
Image caption generation is becoming one of the hot research topics and attracts various researchers. It is a complex process because it utilizes both NLP (natural language processing) and computer vision approaches for generating the tasks. A range of strategies are available for image captioning that connect the visual material with everyday language, such as explaining images with textual descriptions. Pre-trained classification networks like CNN and RNN-based neural network models are used in the literature to encrypt visual data. Even though various literature works have analyzed outstanding image caption techniques, they still lack in providing better performance for diverse databases. To overcome such issues, this research work presents an automated optimization deep learning model for image caption generation. Initially, the input image is pre-processed, and then the encoder decoder-based structure is utilized for extracting the visual features and caption generation. On the encoder side, the pre-trained ResNet 101 (residual network) is used to extract the visual features, and the SA- Bi-LSTM (self-attention with bi-directional Long Short-Term Memory) is used to generate the caption on the decoder side. In addition, an optimization model CA (Chimp algorithm) is used to improve detection performance in caption generation. The proposed encoder-decoder model is tested on benchmark datasets like Flickr8k, Flickr30k and COCO. Further, this model attained better BLEU and ribes scores of 0.8595 and 0.3531 on the Flickr8k dataset. Thus, the proposed SA-BiLSTM model achieved a significant performance in image caption generation.
Keyword:
Chimp Algorithm
Deep Learning Models
Decoder
Encoder
Image Caption Generation
Visual Features
期刊
IF:
3
论文数:
1.9W
被引数:
3.2W
机构
引用论文
Image caption generation using Visual Attention Prediction and Contextual Spatial Relation Extraction
JOURNAL OF BIG DATA
IF6.4
Deep Learning Approaches Based on Transformer Architectures for Image Captioning Tasks基于Transformer架构的图像字幕任务深度学习方法
IEEE ACCESS
IF3.6
Image captioning model using attention and object features to mimic human image understanding
JOURNAL OF BIG DATA
IF6.4

