返回
An encoder-decoder based framework for hindi image caption generation
DOI:10.1007/s11042-021-11106-5.png)
摘要
En 中文
In recent times, research activity on image caption generation has attracted several researchers. The present work attempt to address the problem of Hindi image caption generation using Hindi Visual genome dataset. Hindi is the official and most spoken language in India. In a linguistically diverse country like India, it is essential to provide a means that can help the people to understand the visual entities in their native languages. In this paper, an encoder-decoder based architecture is proposed where Convolutional Neural Network (CNN) is employed for encoding visual features of an image and stacked Long Short-Term Memory (sLSTM) in combination with both uni-directional LSTM and bi-directional LSTM for generating the captions in Hindi. For encoding the visual feature representation of an image, VGG19 based pre-trained model is used and sLSTM architecture is employed for caption generation at the decoder side. The model is tested over Hindi visual genome dataset to validate the proposed approach's performance and cross-verification is carried out for English captions with Flickr dataset. The experimental results of the proposed approach manifest that the model is qualitatively and quantitatively better than state-of-the-art approaches for Hindi caption generation.
Keyword:
CNN
VGG16
VGG19
Stacked LSTM
Hindi image caption generation
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3
论文数:
2.0W
被引数:
3.2W
机构
引用论文
Cysteine-specific (89)Zr-labeled anti-CD25 IgG allows immuno-PET imaging of interleukin-2 receptor-α on T cell lymphomas半胱氨酸特异性(89)Zr标记的抗CD25 IgG能够实现白细胞介素-2受体-α在T细胞淋巴瘤中的免疫PET成像

