Return
Code-Driven LLM Agent for One-Shot Explanatory Visual Question Answering
Z
D
B
S
C
DOI:10.1145/3785327.png)
Abstract
En 中文
Code-driven Large Language Models (LLMs) integrate both natural and formal languages, enhancing reasoning, precision, and interaction with execution environments, which in turn augments the capabilities of intelligent agents. Recent advancements in code-driven LLMs have proven pivotal for visual tasks, such as Visual Question Answering (VQA), a critical task at the intersection of computer vision and natural language processing. Despite significant progress, interpretability remains a challenge for VQA models, leading to the emergence of multimodal explanations for VQA. In this article, we propose the One-Shot and Training-Free Code-Driven LLM Agent (OneCoLA), a novel framework for Multimodal Explanatory Visual Question Answering (MEVQA). OneCoLA enables LLMs to generate multimodal explanations for the VQA task by utilizing a one-shot prompt to convert input questions into Python programs that model the reasoning process. The framework supports the flexible integration of open-world tools, ensuring adaptability to different problem contexts. During program execution, OneCoLA captures and preserves key execution data to enhance the interpretability of the results. Additionally, through further one-shot prompting, the framework generates multimodal explanations by combining execution outcomes with relevant visual content, providing both textual and visual context. Experimental results compared with state-of-the-art methods demonstrate the effectiveness of OneCoLA in generating accurate and interpretable multimodal explanations without the need for extensive training.
Keywords:
Large Language Models
Explanatory Visual Question Answering
Few-Shot Learning
Journal
IF:
6
Papers:
2.0K
Citations:
5.4K
