arrow
Return

Multilayer Dense Attention Model for Image Caption

delete2019-01-01
delete25
delete
OA
AI
E
Eric Ke Wang
X
Xun Zhang
王凡 (Fan Wang)
T
Tsu‐Yang Wu
C
Chien‐Ming Chen *
DOI:10.1109/ACCESS.2019.2917771delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
The image caption is a technology that enables us to understand the contents and generate descriptive text, of images using machines. With the development of deep learning, means of using it to understand image content and generate descriptive text has become a hot research topic. This paper proposes a multilayer dense attention model for image caption. A faster recurrent convolutional neural networks (Faster R-CNN) is employed to extract image features as the coding layer, the long short-term memory (LSTM)-attend is used to decode the multilayer dense attention model, and the description text is generated. The model parameters are optimized using strategy gradient optimization in reinforcement learning. Use of dense attention mechanisms in the coding layer can effectively avoid the interference of non-salient information and selectively output the corresponding description text for the decoding process. The experimental results in the field of general images validate the model's good ability to understand images and generating text.
Keywords:
Attention
image caption
LSTM
RCNN
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66