1
Return

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

delete2026-05-05
delete0
PRE
AI
J
Jiacheng Li
P
Ping Wei
W
Wenjuan Han
S
Song-Chun Zhu
L
Lifeng Fan
DOI:10.1109/tpami.2026.3690561delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions—often termed the “dark matter” of social intelligence. To bridge the gap between visual observation and intent reasoning, we introduce a novel task, <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">IntentQA</b>, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a “Contrast Performance Decline” metric. We propose the <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">X-CaVIR</b>(eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of “Cognitive Context” to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the lack of transparency in traditional models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets.
Keywords:
Intent understanding
social intelligence
video question answering
visual reasoning
context

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

B
Beijing Jiaotong University
Scholars:
2.1W
Papers: 1.7W
Citations: 1.2W
X
xi'an jiaotong university
Scholars:
8.9W
Papers: 6.5W
Citations: 75
B
Cited Papers

Cited Papers

Citing Papers

Citing Papers