arrow
返回

Prompting large language model with context and pre-answer for knowledge-based VQA

delete2024-07-01
delete3
PRE
AI
Z
Zhongjian Hu
杨鹏 封面图
杨鹏 (Peng Yang) *
Z
Zijian Bai
DOI:10.1016/j.patcog.2024.110399delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Existing studies apply Large Language Model (LLM) to knowledge -based Visual Question Answering (VQA) with encouraging results. Due to the insufficient input information, the previous methods still have shortcomings in constructing the prompt for LLM, and cannot fully activate the capacity of LLM. In addition, previous works adopt GPT-3 for inference, which has expensive costs. In this paper, we propose PCPA: a framework that Prompts LLM with Context and Pre -Answer for VQA. Specifically, we adopt a vanilla VQA model to generate in -context examples and candidate answers, and add a pre -answer selection layer to generate preanswers. We integrate in -context examples and pre -answers into the prompt to inspire the LLM. In addition, we choose LLaMA instead of GPT-3, which is an open and free model. We build a small dataset to fine-tune the LLM. Compared to existing baselines, the PCPA improves accuracy by more than 2.1 and 1.5 on OK-VQA and A-OKVQA, respectively.
Keyword:
Visual question answering
Large language model
Knowledge-based VQA
Fine-tuning
In-context learning

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

S
southeast university - china
学者数:
5.3W
论文数: 4.9W
被引数: 57
引用论文

引用论文

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
err分享
err收藏
Learning visual question answering on controlled semantic noisy labels在受控语义噪声标签上学习视觉问答
err2023-06-01
err16
PREAI
errZhang, Haonan; Zeng, Pengpeng; Hu, Yuxuan; Qian, Jin; Song, Jingkuan; Gao, Lianli
err分享
err收藏
Visual question answering from another perspective: CLEVR mental rotation tests *
err2023-04-01
err4
errOAAI
errBeckham, Christopher; Weiss, Martin; Golemo, Florian; Honari, Sina; Nowrouzezahrai, Derek; Pal, Christopher
err分享
err收藏
err分享
err收藏
Hierarchical multimodal transformers for Multipage DocVQA
err2023-12-01
err10
errOAAI
errTito, Ruben; Karatzas, Dimosthenis; Valveny, Ernest
err分享
err收藏
学者 查看更多内容