1
Return

Code-Driven LLM Agent for One-Shot Explanatory Visual Question Answering

delete2026-03-01
delete0
PRE
AI
Z
Zhou, Zuyi
D
Dizhan Xue
B
Baoyuan Qi
S
Shengsheng Qian *
C
Changsheng Xu
DOI:10.1145/3785327delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Code-driven Large Language Models (LLMs) integrate both natural and formal languages, enhancing reasoning, precision, and interaction with execution environments, which in turn augments the capabilities of intelligent agents. Recent advancements in code-driven LLMs have proven pivotal for visual tasks, such as Visual Question Answering (VQA), a critical task at the intersection of computer vision and natural language processing. Despite significant progress, interpretability remains a challenge for VQA models, leading to the emergence of multimodal explanations for VQA. In this article, we propose the One-Shot and Training-Free Code-Driven LLM Agent (OneCoLA), a novel framework for Multimodal Explanatory Visual Question Answering (MEVQA). OneCoLA enables LLMs to generate multimodal explanations for the VQA task by utilizing a one-shot prompt to convert input questions into Python programs that model the reasoning process. The framework supports the flexible integration of open-world tools, ensuring adaptability to different problem contexts. During program execution, OneCoLA captures and preserves key execution data to enhance the interpretability of the results. Additionally, through further one-shot prompting, the framework generates multimodal explanations by combining execution outcomes with relevant visual content, providing both textual and visual context. Experimental results compared with state-of-the-art methods demonstrate the effectiveness of OneCoLA in generating accurate and interpretable multimodal explanations without the need for extensive training.
Keywords:
Large Language Models
Explanatory Visual Question Answering
Few-Shot Learning

Journal

ACM Transactions on Multimedia Computing Communications and Applications cover
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
Papers:
2.0K
Citations:
5.4K

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 74
C
chinese academy of sciences
Scholars:
54.9W
Papers: 44.5W
Citations: 703
Cited Papers

Cited Papers

Citing Papers

Citing Papers