arrow
Return

Data-And Knowledge-Driven Visual Abductive Reasoning

delete2025-09-23
delete0
PRE
AI
L
Liang Chen
W
Wenguan Wang
L
Ling Chen
Y
Yi Yang
DOI:10.1109/TPAMI.2025.3613712delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Abductive reasoning seeks the likeliest possible explanation for partial observations. Although being frequently employed in human daily reasoning, abduction is rarely explored in computer vision literature. In this article, we propose a new task, Visual Abductive Reasoning (VAR), that underpins the machine intelligence study of abductive reasoning in everyday visual situations. Given an incomplete set of visual events, AI systems are required to not only describe what is observed, but also infer the hypothesis that can best explain the observed premise. We create the first large-scale VAR dataset, which contains a total of 9K examples. We further devise a transformer-based VAR model – <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Reasoner</small>v2 – for knowledge-driven, causal-and-cascaded reasoning. <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Reasoner</small>v2 first adopts a contextualized directional position embedding strategy in the encoder, to capture the causal-related temporal structure of the observations, and yield discriminative representations for the premises and hypotheses. Then, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Reasoner</small>v2 extracts condensed causal knowledge from external knowledge bases, for reasoning beyond observation. Finally, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Reasoner</small>v2 cascades multiple decoders so as to generate and progressively refine the premise and hypothesis sentences. The prediction scores of the sentences are used to guide cross-sentence information flow in the cascaded reasoning procedure. Our VAR benchmarking results show that <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Reasoner</small>v2 surpasses many famous video-language models, while still being far behind human performance.
Keywords:
Visual abductive reasoning
causal knowledge
multimodal transformer

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

U
university of technology sydney
Scholars:
1.6W
Papers: 2.0W
Citations: 25
Z
zhejiang university
Scholars:
17.4W
Papers: 12.0W
Citations: 152