arrow
返回

Explicitly diverse visual question generation

delete2025-04-01
delete0
PRE
AI
J
Jiayuan Xie
Z
Zheng, Jiasheng
W
Wenhao Fang
蔡毅 封面图
蔡毅 (Yi Cai) *
Prof. LI Qing 封面图
Prof. LI Qing (Qing Li)
DOI:10.1016/j.neunet.2024.107002delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual question generation involves the generation of meaningful questions about an image. Although we have made significant progress in automatically generating a single high-quality question related to an image, existing methods often ignore the diversity and interpretability of generated questions, which are important for various daily tasks that require clear question sources. In this paper, we propose an explicitly diverse visual question generation model that aims to generate diverse questions based on interpretable question sources. To explicitly perform question generation, our model first extracts the scene graph from the image using the unbiased scene graph generation method, where questions generated based on the scene graphs have interpretable question sources. To ensure the diversity of generated questions, our model selects different subgraphs from the scene graph as question sources. Specifically, we employ a subgraph selector to learn how humans select multiple subgraphs that are suitable for question generation. Finally, our model generates diverse questions based on different selected subgraphs. Extensive experiments on the VQA v2.0 and COCO-QA datasets show that the proposed model outperforms the baselines and is able to interpretably generate diverse questions.
Keyword:
Multimodal
Diverse visual question generation
Interpretable text generation
Unbiased scene graph generation

期刊

Neural Networks 封面图
Neural Networks
IF:
6.3
论文数:
8.2K
被引数:
3.0W

机构

H
hong kong polytechnic university
学者数:
3.0W
论文数: 4.1W
被引数: 921
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
S
south china university of technology
学者数:
6.8W
论文数: 5.1W
被引数: 85
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Publisher Correction: An automated Raman-based platform for the sorting of live cells by functional properties
err2019-04-12
err0
errOAAI
errKang Soo Lee; Márton Palatinszky; Fátima C. Pereira; Jen Nguyen; Vicente I. Fernandez; Anna J. Mueller; Filippo Menolascina; Holger Daims; David Berry; Michael Wagner; Roman Stocker
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
学者 查看更多内容