arrow
返回

Cross modification attention-based deliberation model for image captioning

delete2022-07-05
delete4
PRE
AI
Z
Zheng Lian *
Y
Yanan Zhang
H
Haichang Li
王芮 封面图
王芮 (Rui Wang)
X
Xiaohui Hu
DOI:10.1007/s10489-022-03845-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The two-pass decoding framework has been proved to considerably improve the performance of image captioning models. However, most of the existing two-pass models involve the coarse captions in assisting the refining process by simply using a conventional attention module. Such an insufficient interaction cannot provide satisfactory support for reproducing higher-quality image descriptions. In this paper, we propose a novel Cross Modification Attention (CMA) module to exploit the complementarity of images and the corresponding coarse captions to supply more reliable features for refinement. Specifically, our CMA extends the conventional attention mechanisms with a hierarchical gating network, which mutually modifies the attended vectors of both visual and linguistic modalities. Thus, it can make the visual semantic representation more unambiguous and filter out misleading information from the coarse captions. To cooperate with CMA in feature interaction, we further explore a general two-pass decoding framework, where the drafting and the deliberation model share only the image encoders rather than the whole drafting network as previous methods. Our framework provides visual features tightly coupling both decoding processes, and ensures the efficient joint optimization of the two-pass models. Moreover, we consider the coarse captions as a baseline when optimizing the deliberation model and employ a potential-oriented reward shaping strategy for reinforcement learning to pertinently improve the quality of refinement. Experiments on Flickr30K and MS COCO datasets demonstrate that our Cross Modification Attention-based Deliberation Model (CMA-DM) obtains significant improvements over single-pass decoding baselines and achieves competitive performance on MS COCO online test server.
Keyword:
Image captioning
Two-pass decoding
Deliberation
Attention mechanism
Reinforcement learning

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

3G structure for image caption generation
err2019-02-01
err29
errOAAI
errYuan, Aihong; Li, Xuelong; Lu, Xiaoqiang
err分享
err收藏
Remote sensing of fish-processing in the Sundarbans Reserve Forest, Bangladesh: an insight into the modern slavery-environment nexus in the coastal fringe
err2020-09-17
err0
errOAAI
errBethany Jackson; Doreen S. Boyd; Christopher D. Ives; Jessica L. Decker Sparks; Giles M. Foody; Stuart Marsh; Kevin Bales
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Labral Augmentation with Native Tissue Preservation with a 7.5-Year Follow-up
err2018-03-28
err0
PREAI
errJonathan A. Godin; Lorenzo Fagotti; Karen K. Briggs; Marc J. Philippon
err分享
err收藏
Liquid structure of the alkaline-earth metals
err1993-06-01
err0
PREAI
errL. E. González; A. Meyer; M. P. Iñiguez; D. J. González; M. Silbert
err分享
err收藏
学者 查看更多内容