arrow
Return

QC 2-VQG: Question context complement for visual question generation

delete2025-08-20
delete0
PRE
AI
Y
Ying Zhang *
X
Xubo Liu
Z
Ziyu Lu
W
Wenya Guo
X
Xumeng Liu
R
Ruxue Yan
DOI:10.1016/j.knosys.2025.114306delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual Question Generation (VQG) is a critical vision-language understanding task that involves generating human-like questions from given images and associated textual information. Most of the existing works are answer-aware and focus on modeling the complex relationship between the answer and its relevant object regions. However, we observe that these approaches disregard question context (e.g., the image description and visual entities) which is indispensable for a large number of questions, making it difficult to generate the desired questions. To address this issue, we present a novel strategy to generate questions by supplementing image-related question context. The key motivation is that the question context can bridge the task gap between visual understanding and question generation. Thus, we propose QC 2-VQG which can automatically capture the visual information and convert it to the textual question context for the answer-aware QG. Extensive experiments on two widely used datasets demonstrate that QC 2-VQG outperforms SOTA methods across various evaluation metrics, highlighting its effectiveness in generating high-quality, contextually grounded questions.
Keywords:
Visual Question Generation
Question Context
Answer-aware
Image Description
Vision-Language Understanding

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

No organization information available