Return
Context-aware transformer for image captioning
DOI:10.1016/j.neucom.2023.126440.png)
Abstract
En 中文
Recently, image captioning models have made remarkable progress by introducing transformer architecture, which utilizes self-attention to explore intra- and inter-modal interactions. However, most existing methods only consider region-level characteristic during the attention weight calculation and ignore the image-level information. This seriously hinders the whole model from understanding the scene content. In this paper, we propose a Context-Aware Transformer (CATNet) with two novel designs, namely Context Augmented Attention (CAA) and Dual Way Controller (DWC). Concretely, CAA in encoder enables the extraction of more comprehensive visual representation through modeling the communications between multi-level visual features. DWC in decoder is used to enhance the fusion between visual features and language representation through utilizing complementarity of global context and local regions. Extensive experiments conducted on MSCOCO dataset show that the proposed CATNet has achieved state-of-the-art performance on both Karpathy test set and online test. (C) 2023 Published by Elsevier B.V.
Keywords:
Image captioning
Transformer
Attention mechanism
Global context
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

