Return
QC 2-VQG: Question context complement for visual question generation
DOI:10.1016/j.knosys.2025.114306.png)
Abstract
En 中文
Visual Question Generation (VQG) is a critical vision-language understanding task that involves generating human-like questions from given images and associated textual information. Most of the existing works are answer-aware and focus on modeling the complex relationship between the answer and its relevant object regions. However, we observe that these approaches disregard question context (e.g., the image description and visual entities) which is indispensable for a large number of questions, making it difficult to generate the desired questions. To address this issue, we present a novel strategy to generate questions by supplementing image-related question context. The key motivation is that the question context can bridge the task gap between visual understanding and question generation. Thus, we propose QC 2-VQG which can automatically capture the visual information and convert it to the textual question context for the answer-aware QG. Extensive experiments on two widely used datasets demonstrate that QC 2-VQG outperforms SOTA methods across various evaluation metrics, highlighting its effectiveness in generating high-quality, contextually grounded questions.
Keywords:
Visual Question Generation
Question Context
Answer-aware
Image Description
Vision-Language Understanding
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
Organization
No organization information available

