arrow
返回

Interaction augmented transformer with decoupled decoding for video captioning q

delete2022-07-01
delete9
PRE
AI
T
Tao Jin *
Z
Zhou Zhao
王鹏 封面图
王鹏 (Peng Wang)
J
Jun Yu
Wu Fei 封面图
Wu Fei (Fei Wu)
DOI:10.1016/j.neucom.2022.03.065delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Transformer-based architectures achieve competitive performances in video captioning. However, their applicability still has many issues: (1) Existing methods only consider the correlation of query and key modalities when calculating the attention weights, and ignore their interaction with other modalities. (2) Deep stacked cross-modal encoding blocks make the different modalities assimilative and lose their preliminary discriminative properties. (3) The decoder usually employs the output of the last encoding block, which is not a comprehensive representation. Based on these concerns, we propose a novel method called Interaction Augmented Transformer (IAT) with discriminative encoding and decoupled decoding for video captioning. Concretely, by concatenating [CLS] tokens to multimodal features, we perform reconstructive contrastive constraints for the encoded results. Based on the conclusive information carried by these tokens, we first introduce the global-gated interaction into multi-head attention, where the conclusive information mentioned above is transformed into multiple interaction augmented functions. Additionally, the dot-product operation is replaced by the tucker-fused operation to better capture the query-to-key correlation. Furthermore, we employ fine-grained layer-wise decoding for multi-layer multi-modal features from the encoder with decoupled strategy. We conduct extensive quantitative, qualitative, and ablation experiments on the benchmark datasets and the experimental results show that IAT outperforms the state-of-the-art methods under most of the metrics. (c) 2022 Published by Elsevier B.V.
Keyword:
Multi-modal video captioning
Transformer
Interaction augmentation
Decoupled decoding

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

H
Hangzhou Dianzi University
学者数:
1.3W
论文数: 9.6K
被引数: 7.5K
N
Northwestern Polytechnical University
学者数:
4.6W
论文数: 3.7W
被引数: 5.3W
Z
zhejiang university
学者数:
17.7W
论文数: 12.1W
被引数: 152
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Remote sensing of fish-processing in the Sundarbans Reserve Forest, Bangladesh: an insight into the modern slavery-environment nexus in the coastal fringe
err2020-09-17
err0
errOAAI
errBethany Jackson; Doreen S. Boyd; Christopher D. Ives; Jessica L. Decker Sparks; Giles M. Foody; Stuart Marsh; Kevin Bales
err分享
err收藏
Structural phase transitions in aluminium above 320 GPa
err2018-09-28
err0
errOAAI
errGuillaume Fiquet; Chandrabhas Narayana; Christophe Bellin; Abhay Shukla; Imène Estève; Art L. Ruoff; Gaston Garbarino; Mohamed Mezouar
err分享
err收藏
Re-Caption: Saliency-Enhanced Image Captioning Through Two-Phase Learning
err2020-01-01
err52
PREAI
errZhou, Lian; Zhang, Yuejie; Jiang, Yu-Gang; Zhang, Tao; Fan, Weiguo
err分享
err收藏
学者 查看更多内容