Return
Visual Question Explainable Reasoning on Hypothesis Agent Interaction with Scene
DOI:10.1016/j.sigpro.2025.110177.png)
Abstract
En 中文
• We introduce the explainable interaction visual question-answering task, which requires answering questions that include reasoning about human activity beyond the given image and explaining how to obtain the answers which makes the answer more trustworthy. • We build a VQER task for evaluating human action-related questions outside the image. • We propose a explaining and answering model by fusing heterogeneous information.
Keywords:
Visual Question Answering
Multimodal processing
Knowledge graph
Scene understanding
Trustworthy reasoning
Journal
IF:
3.6
Papers:
9.9K
Citations:
1.7W

