返回
Optimal transformers based image captioning using beam search
DOI:10.1007/s11042-023-17359-6.png)
摘要
En 中文
Image Captioning is the process of generating textual descriptions of given images. It encompasses two major fields of deep learning, computer vision, and natural language processing. This paper presents an Image Captioning model which uses the Convolution Neural Network (CNN) model for feature extraction and a transformer architecture for the generation of sequences from these feature vectors. For feature extraction, this paper uses different CNN architectures like Xception, InceptionV3, ResNet50V2, VGG19, DenseNet201, ResNet152V2, EfficientNetV2B3, EfficientNetV2B0. The proposed method takes advantage of the transformer model for faster processing, and Beam search is implemented to get the top N most probable sequences for each image. The architecture is trained on Flickr8k dataset and the model outperforms the existing methods. The proposed model achieves a BLEU_4 score of 0.2184 on the Flickr8k dataset.
Keyword:
Image captioning
Deep learning
Attention
Computer vision
Sequence models
期刊
IF:
3
论文数:
2.0W
被引数:
3.2W
机构
暂无机构信息
引用论文
Phrase-based image caption generator with hierarchical LSTM network基于短语的分层LSTM网络图像字幕生成器
NEUROCOMPUTING
IF6.5

