返回
Prompting large language model with context and pre-answer for knowledge-based VQA
DOI:10.1016/j.patcog.2024.110399.png)
摘要
En 中文
Existing studies apply Large Language Model (LLM) to knowledge -based Visual Question Answering (VQA) with encouraging results. Due to the insufficient input information, the previous methods still have shortcomings in constructing the prompt for LLM, and cannot fully activate the capacity of LLM. In addition, previous works adopt GPT-3 for inference, which has expensive costs. In this paper, we propose PCPA: a framework that Prompts LLM with Context and Pre -Answer for VQA. Specifically, we adopt a vanilla VQA model to generate in -context examples and candidate answers, and add a pre -answer selection layer to generate preanswers. We integrate in -context examples and pre -answers into the prompt to inspire the LLM. In addition, we choose LLaMA instead of GPT-3, which is an open and free model. We build a small dataset to fine-tune the LLM. Compared to existing baselines, the PCPA improves accuracy by more than 2.1 and 1.5 on OK-VQA and A-OKVQA, respectively.
Keyword:
Visual question answering
Large language model
Knowledge-based VQA
Fine-tuning
In-context learning
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
Learning visual question answering on controlled semantic noisy labels在受控语义噪声标签上学习视觉问答
PATTERN RECOGNITION
IF7.6
Visual question answering from another perspective: CLEVR mental rotation tests *
PATTERN RECOGNITION
IF7.6
Image captioning for effective use of language models in knowledge-based visual question answering在基于知识的视觉问答中有效使用语言模型的图像字幕

