arrow
Return

PSNet: position-shift alignment network for image caption

delete2023-11-27
delete3
PRE
AI
L
Lixia Xue
A
Awen Zhang
R
Ronggui Wang
杨娟 (Juan Yang) *
DOI:10.1007/s13735-023-00307-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recently, Transformer-based models have gained increasing popularity in the field of image captioning. The global attention mechanism of the Transformer facilitates the integration of region and grid features, leading to a significant improvement in accuracy. However, combining two features through direct fusion may lead to inevitable semantic noise, which is caused by non-synergistic issue of the region and grid features; meanwhile, the additional detector to extract region features also decrease the efficiency of the model. In this paper, we introduce a novel position-shift alignment network (PSNet) to exploit the advantages of the two features. Concretely, we embed a simple detector DETR into the model and extracted region features based on grid features to improve model efficiency. Moreover, we propose a P-shift alignment module to address semantic noise caused by non-synergistic issue of the region and grid features. To validate our model, we conduct extensive experiments and visualization on the MS-COCO dataset, and results show that PSNet is qualitatively competitive with existing models under comparable experimental conditions.
Keywords:
Image caption
Grid features
Region features
Transformer

Journal

International Journal of Multimedia Information Retrieval cover
International Journal of Multimedia Information Retrieval
IF:
2.9
Papers:
273
Citations:
866

Organization

H
hefei university of technology
Scholars:
2.5W
Papers: 1.7W
Citations: 35