arrow
返回

Instance-sequence reasoning for video question answering

delete2022-04-02
delete17
PRE
AI
R
Rui Liu
Y
Yahong Han *
DOI:10.1007/s11704-021-1248-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video question answering (Video QA) involves a thorough understanding of video content and question language, as well as the grounding of the textual semantic to the visual content of videos. Thus, to answer the questions more accurately, not only the semantic entity should be associated with certain visual instance in video frames, but also the action or event in the question should be localized to a corresponding temporal slot. It turns out to be a more challenging task that requires the ability of conducting reasoning with correlations between instances along temporal frames. In this paper, we propose an instance-sequence reasoning network for video question answering with instance grounding and temporal localization. In our model, both visual instances and textual representations are firstly embedded into graph nodes, which benefits the integration of intra- and inter-modality. Then, we propose graph causal convolution (GCC) on graph-structured sequence with a large receptive field to capture more causal connections, which is vital for visual grounding and instance-sequence reasoning. Finally, we evaluate our model on TVQA+ dataset, which contains the groundtruth of instance grounding and temporal localization, three other Video QA datasets and three multimodal language processing datasets. Extensive experiments demonstrate the effectiveness and generalization of the proposed method. Specifically, our method outperforms the state-of-the-art methods on these benchmarks.
Keyword:
video question answering
instance grounding
graph causal convolution

期刊

Frontiers of Computer Science 封面图
Frontiers of Computer Science
IF:
4.6
论文数:
1.6K
被引数:
2.8K

机构

T
tianjin university
学者数:
8.0W
论文数: 5.8W
被引数: 88
引用论文

引用论文

Adaptation to novel environments during crop diversification
err2020-08-01
err0
PREAI
errGaia Cortinovis; Valerio Di Vittori; Elisa Bellucci; Elena Bitocchi; Roberto Papa
err分享
err收藏
The Utility of Combining the IAD and SES Frameworks
err2019-05-07
err0
errOAAI
errDaniel H. Cole; Graham Epstein; Michael D. McGinnis
err分享
err收藏
Sequential Video VLAD: Training the Aggregation Locally and Temporally
err2018-10-01
err91
PREAI
errXu, Youjiang; Han, Yahong; Hong, Richang; Tian, Qi
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Generation of spin cat states in an engineered Dicke model
err2021-11-29
err0
errOAAI
errCaspar Groiseau; Stuart J. Masson; Scott Parkins
err分享
err收藏
学者 查看更多内容