arrow
Return

A visual persistence model for image captioning

delete2022-01-01
delete18
PRE
AI
Y
Yiyu Wang
J
Jungang Xu *
Y
Yingfei Sun
DOI:10.1016/j.neucom.2021.10.014delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Object-level features from Faster R-CNN and attention mechanism have been used extensively in image captioning based on Encoder-Decoder frameworks. However, most existing methods feed the average pooling of object features as the global representation to the captioning model and recalculate the attention weights of object regions when generating a new word without considering the visual persistence like humans. In this paper, we respectively build Visual Persistence modules in encoder and decoder: The visual persistence module in encoder seeks the core object features to replace the image global representation; the visual persistence module in decoder evaluates the correlation between previous attention results and current attention results, and fuses them as the final attended feature to generate a new word. The experimental results on MSCOCO validate the effectiveness and competitiveness of our Visual Persistence Model (VPNet). Remarkably, VPNet also achieves competitive scores in most metrics on MSCOCO online test server compared to the existing state-of-the-art methods. (c) 2021 Elsevier B.V. All rights reserved.
Keywords:
Image captioning
Attention mechanism
Visual persistence

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 75
C
chinese academy of sciences
Scholars:
56.3W
Papers: 44.8W
Citations: 704