arrow
Return

Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering

delete2026-01-12
delete0
PRE
AI
J
Jiangtong Li
Z
Zhaohe Liao
F
Fengshun Xiao
T
Tianjiao Li
Q
Qiang Zhang
H
Haohua Zhao
L
Li Niu
陈光 cover
陈光 (Guang Chen)
张丽清 (Liqing Zhang)
C
Changjun Jiang
DOI:10.1109/TPAMI.2026.3650864delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula>), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula> consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula> Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula> framework on QPVA<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula> Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts.
Keywords:
Multi-modal large language model
multimodal reasoning framework
multi-modal benchmark
compositional reasoning
video question-answering

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

S
shanghai jiao tong university
Scholars:
15.5W
Papers: 11.6W
Citations: 159
T
tongji university
Scholars:
7.7W
Papers: 5.9W
Citations: 98
B
bilibili inc
Scholars:
7
Papers: 2
Citations: 0
researcher View more organizations