Return
Automatic image captioning system using a deep learning approach
DOI:10.1007/s00500-023-08544-8.png)
Abstract
En 中文
This paper's residual network is tailored to increase the high-quality image caption generation ability. The captioning is exploited using the relevant content with high-quality interpretation. The research develops a Residual Attention Generative Adversarial Network (RAGAN) and uses attention-based residual learning in Generative Adversarial Network (GAN) to improve the diversity and fidelity of the generated image captions. The RAGAN exploits the words based on the feature maps faster to generate high-quality captions. The RAGAN improves the diversity of captions generated and increases the language metrics scores. The generator is designed as an encoder-decoder mechanism that operates in an unsupervised manner. The residual learning is adopted between the encoder and decoder network. The discriminator is connected to a language evaluator unit, which provides feed-forward to the generator and discriminator to either positively or negatively influence the image captioning process. The experiments show that the proposed RAGAN performs better than the state-of-the-art GAN models.
Keywords:
Image captioning
Deep learning
Generative adversarial network
Residual learning
Journal
IF:
2.5
Papers:
1.0W
Citations:
2.1W

