arrow
返回

Multimodal graph neural network for video procedural captioning

delete2022-06-01
delete6
PRE
AI
L
Lei Ji *
R
Rong-Cheng Tu
K
Kevin Lin
王丽娟 封面图
王丽娟 (Lijuan Wang)
N
Nan Duan
DOI:10.1016/j.neucom.2022.02.062delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video procedural captioning aims to generate detailed descriptive captions for all steps in a long instructional video. The peculiarity of this problem is the procedural dependency between the events to generate consistent captions among the video. However, existing video (dense) captioning methods only consider intra-event or sequential inter-event context and are hard to model the non-sequential context dependency between events. In this paper, inspired by the recent success of graph neural networks in capturing the relations for structured data, we propose a novel Multimodal Graph Neural Network (MGNN) for dense video procedural captioning in capturing the procedural structure between events. Specifically, we construct temporal sequential graph and semantic non-sequential graph for a multi modal heterogeneous graph. Moreover, we adopt the graph neural network to enhance the visual and text features, and fuse both features for further caption generation. Extensive experiments demonstrate the proposed MGNN is effective in generating coherent captions on both the Youcook2 and Activitynet Captions benchmark.(c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Multimodal video captioning
Graph neural network

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
M
Microsoft Research Asia
学者数:
421
论文数: 407
被引数: 2
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

Publish or Perish
err2005-12-01
err0
PREAI
errMark De Rond; Alan N. Miller
err分享
err收藏