arrow
Return

Attribute-driven image captioning via soft-switch pointer

delete2021-12-01
delete10
PRE
AI
Y
Yujie Zhou
J
Jiefeng Long
S
Suping Xu
L
Lin Shang *
DOI:10.1016/j.patrec.2021.08.021delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual attributes detection provides rich semantic concepts for image captioning. Some previous methods attempt to directly encode the attributes into vectors and generate the corresponding captions, which ignore the correlations between the image regions and attributes. In this paper, we consider to bridge the gap between visual features and detected attributes: first to look at a specific region of the image and second to decide which attribute to attend to. We propose an attribute-driven image captioning approach consisting of two parts: the visual positioning part and the attribute selection part. Specifically, we introduce the pointer-generator network into the second part of our model as a soft-switch, which determines whether to generate a word through the hidden state or point to a detected attribute at each decoding step. Qualitative and Quantitative experiments show that our model can improve the coverage of key visual attributes and significantly boost the overall performance. (c) 2021 Published by Elsevier B.V.
Keywords:
Image captioning
Visual attributes detection
Attention
Pointing mechanism
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.8K
Citations:
1.6W

Organization

N
nanjing university
Scholars:
7.7W
Papers: 5.6W
Citations: 87