Return
Explainable Knowledge reasoning via thought chains for knowledge-based visual question answering
DOI:10.1016/j.ipm.2024.103726.png)
Abstract
En 中文
Knowledge -based visual question answering (KBVQA) requires external knowledge for crossmodal scene understanding and reasoning. Activating the reasoning capability of language models, such as Large Language Models (LLMs), and generating explanations for consistency presents a formidable challenge. We claim that incorporating intermediate reasoning chains enriched with multimodal knowledge is essential for KBVQA tasks. To achieve this, we present Multimodal Knowledge Reasoning via Chain -of -Thought (MuKCoT) for KBVQA. The idea is to leverage the chain -of -thought capability of LLMs with vision -grounded knowledge to bridge the explanation gap for the answers. Our MuKCoT generates auto -labeling reasoning chains via LLMs prompting and train smaller Vision -and -language models to perform CoT reasoning for KBVQA tasks. Our approach releases the annotated knowledge resources and explanations. The MuKCoT algorithm, which has less than 1 billion parameters, does 6.6% better than previous state-of-the-art approaches on the knowledge -based VQA dataset OK-VQA and 1.9% better on the direct -answer task A-OKVQA.
Keywords:
Knowledge-based visual question answering
Knowledge reasoning
Chain of thought
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W
Organization
No organization information available

