arrow
Return

Image captioning with transformer and knowledge graph

delete2021-03-01
delete63
PRE
AI
张宇 (Yu Zhang) *
X
Xinyu Shi
米思娅 cover
米思娅 (Siya Mi)
X
Xu Yang
DOI:10.1016/j.patrec.2020.12.020delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Transformer model has achieved very good results in machine translation tasks. In this paper, we adopt the Transformer model for the image captioning task. To promote the performance of image captioning, we improve the Transformer model from two aspects. First, we augment the maximum likelihood estimation (MLE) with an extra Kullback-Leibler (KL) divergence term to distinguish the difference between incorrect predictions. Second, we introduce a method to help the Transformer model generate captions by leveraging the knowledge graph. Experiments on benchmark datasets demonstrate the effectiveness of our method. (c) 2021 Elsevier B.V. All rights reserved.
Keywords:
Image captioning
Transformer
Knowledge graph
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.9K
Citations:
1.6W

Organization

N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
S
southeast university - china
Scholars:
5.3W
Papers: 4.9W
Citations: 57