arrow
返回

Diverse Visual Question Generation Based on Multiple Objects Selection

delete2024-03-08
delete0
PRE
AI
W
Wenhao Fang
J
Jiayuan Xie
H
Hongfei Liu
J
Jiali Chen
蔡毅 封面图
蔡毅 (Yi Cai) *
DOI:10.1145/3640014delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual question generation task aims at generating high-quality questions about a given image. To make this tak applicable to various scenarios, e.g., the growing demand for exams, it is important to generate diverse questions. The existing methods for this task control diverse question generation based on different question types, e.g., what and when. Although different question types lead to description diversity, they cannot guarantee semantic diversity when asking the same objects. Research in the field of psychology shows that humans pay attention to different objects in an image based on their preferences, which is beneficial to constructing semantically diverse questions. According to the research, we propose a multi-selector visual question generation (MS-VQG) model that aims to focus on different objects to generate diverse questions. Specifically, our MS-VQG model employs multiple selectors to imitate different humans to select different objects in a given image. Based on these different selected objects, our MS-VQG model can generate diverse questions corresponding to each selector. Extensive experiments on two datasets show that our proposed model outperforms the baselines in generating diverse questions.
Keyword:
Multimodal
visual question generation
mixture of experts

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

S
south china university of technology
学者数:
6.8W
论文数: 5.1W
被引数: 85
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
MODAL SPLIT ANALYSIS BY BEST-WORST METHOD AND MULTINOMINAL LOGIT MODEL
err2023-03-01
err0
errOAAI
errMichal CINGEL; Marek DRLICIAK; Jan CELKO; Katarína ZABOVSKA
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
学者 查看更多内容