arrow
返回

Relation-aware attention for video captioning via graph learning

delete2023-04-01
delete16
PRE
AI
Y
Yunbin Tu
C
Chang Zhou
郭军军 封面图
郭军军 (Junjun Guo)
H
Huafeng Li
高盛祥 封面图
高盛祥 (Shengxiang Gao)
Z
Zhengtao Yu *
DOI:10.1016/j.patcog.2022.109204delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video captioning often uses an attentive encoder-decoder as the baseline model. However, the conven-tional attention mechanism still remains two problems. First, the attended visual feature is often irrele-vant to the target word state, because the attention process only uses the unidirectional flow from vision to linguistics, while lacking the reverse flow. Second, each attention result is independent, because it is computed only based on the previous word states while not considering the attention information from the past and future. This does not suit the attention habits of human beings. In this paper, we improve the conventional attention mechanism to a relation-aware attention mechanism. To this end, we propose two kinds of graph learning strategies, namely the linguistics-to-vision heterogeneous graph (HTG) and the vision-to-vision homogeneous graph (HMG). The HTG aims to enhance the inter-relation of attention by reversely modeling the relation of each word with respect to every attended visual feature, support-ing proper semantic alignment in between. The HMG aims to enhance the intra-relation of attention by capturing the relations among all of the attended visual features, which can leverage the attention infor-mation from the past and future to guide the current attention process. Extensive experiments on two public datasets show that our proposed method not only significantly improves the baseline model, but also outperforms state-of-the-art methods.(c) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Video captioning
Relation -aware attention
Graph learning

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

T
Tsinghua Shenzhen International Graduate School
学者数:
6.8K
论文数: 4.9K
被引数: 9
引用论文

引用论文

"COME TO THINK OF IT?".
err1986-03-01
err0
PREAI
errJAMES B. STIFF; GERALD R. MILLER
err分享
err收藏
Human-Centric Image Captioning
err2022-06-01
err14
PREAI
errYang, Zuopeng; Wang, Pengbo; Chu, Tianshu; Yang, Jie
err分享
err收藏
Enhancing the alignment between target words and corresponding frames for video captioning
err2021-03-01
err41
PREAI
errTu, Yunbin; Zhou, Chang; Guo, Junjun; Gao, Shengxiang; Yu, Zhengtao
err分享
err收藏
The evaluation of urban spatial quality and utility trade-offs for Post-COVID working preferences: a case study of Hong Kong
err2023-01-27
err0
errOAAI
errQiwei Song; Zhiyi Dou; Waishan Qiu; Wenjing Li; Jingsong Wang; Jeroen van Ameijde; Dan Luo
err分享
err收藏
err分享
err收藏
Potent and stage-specific action of glutathione on the development of goat early embryos in vitro
err2000-01-01
err0
PREAI
errChul-Sang Lee; Deog-Bon Koo; Nanzhu Fang; Yongsun Lee; Sang-Tae Shin; Chang-Sik Park; Kyung-Kwang Lee
err分享
err收藏
STAT: Spatial-Temporal Attention Mechanism for Video CaptioningSTAT: 视频字幕的时空注意机制
err2020-01-01
err301
PREAI
errYan, Chenggang; Tu, Yunbin; Wang, Xingzheng; Zhang, Yongbing; Hao, Xinhong; Zhang, Yongdong; Dai, Qionghai
err分享
err收藏
Learning visual relationship and context-aware attention for image captioning学习视觉关系和上下文感知的图像字幕注意
err2020-02-01
err110
PREAI
errWang, Junbo; Wang, Wei; Wang, Liang; Wang, Zhiyong; Feng, David Dagan; Tan, Tieniu
err分享
err收藏
学者 查看更多内容