arrow
Return

Locate Before Answering: Answer Guided Question Localization for Video Question Answering

delete2024-01-01
delete2
delete
OA
AI
T
Tianwen Qian
R
Ran Cui
J
Jingjing Chen *
P
Pai Peng
X
Xiaowei Guo
Y
Yu–Gang Jiang
DOI:10.1109/TMM.2023.3323878delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video question answering (VideoQA) is an essential task in vision-language understanding, which has attracted numerous research attention recently. Nevertheless, existing works mostly achieve promising performances on short videos of duration within 15 seconds. For VideoQA on minute-level long-term videos, those methods are likely to fail because of lacking the ability to deal with noise and redundancy caused by scene changes and multiple actions in the video. Considering the fact that the question often remains concentrated in a short temporal range, we propose to first locate the question to a segment in the video and then infer the answer using the located segment only. Under this scheme, we propose Locate before Answering (LocAns), a novel approach that integrates a question localization module and an answer prediction module into an end-to-end model. During the training phase, the available answer label not only serves as the supervision signal of the answer prediction module, but also is used to generate pseudo temporal labels for the question localization module. Moreover, we design a decoupled alternative training strategy to update the two modules separately. In the experiments, LocAns achieves state-of-the-art performance on three modern long-term VideoQA datasets, NExT-QA, ActivityNet-QA, and AGQA. Its qualitative examples show the reliable performance of the question localization.
Keywords:
Video question answering
video grounding
cross-modal learning

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

F
fudan university
Scholars:
11.6W
Papers: 7.7W
Citations: 121
A
Australian National University
Scholars:
2.1W
Papers: 2.3W
Citations: 3.9W