arrow
返回

Reciprocal question representation learning network for visual dialog

delete2022-06-16
delete2
PRE
AI
H
Hongwei Zhang *
X
Xiaojie Wang
S
Si Jiang
DOI:10.1007/s10489-022-03795-8delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual dialog task entails an agent to answer a series of questions based on an image and the dialog history. Biases are often observed when the agent over relies on the dialog history. Thus, balanced usage of dialog history is crucial. Existing models usually drop several rounds of dialog history or learn a sparse dialog structure to address the overreliance on such history; however, bias might still exist in the selected dialog history. Therefore, we propose a new model, reciprocal question representation learning network (RQRLN), with less bias from dialog history by learning more accurate history-aware representations of questions. Initially, RQRLN adaptively selects favorable information at the token level from two representations of a question encoded with and without a dialog history. Later, the adaptive question representation is assembled with the corresponding image for the final decoder. We also used a new entropy loss function which further reduces the dialog history-based bias, enabling two different types of representations of the same token to learn interactively. Analysis results on the VisDial v1.0 dataset showed that our proposed model achieved state-of-the-art results in terms of normalized discounted cumulative gain (NDCG). We also demonstrate that our model shows lesser bias and infers more generic answers in comparison with models that use the entire history.
Keyword:
Visual dialog
Attention mechanism
Interactive learning
Multi-model fusion

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

B
beijing university of posts & telecommunications
学者数:
1.4W
论文数: 1.2W
被引数: 9
引用论文

引用论文

A novel intracochlear injection method for rapid drug delivery to vestibular end organs
err2020-07-01
err0
errOAAI
errVishal Raghu; Yugandhar Ramakrishna; Robert F. Burkard; Soroush G. Sadeghi
err分享
err收藏
A mobile lidar system for aerosol and water vapor detection in troposphere with mobile lida
err2016-01-01
err0
PREAI
err吕炜煜 Lv Weiyu; 苑克娥 Yuan Ke′e; 魏 旭 Wei Xu; 刘李辉 Liu Lihui; 王邦新 Wang Bangxin; 吴德成 Wu Decheng; 胡顺星 Hu Shunxing; 王建国 Wang Jianguo; 马振富 Ma Zhenfu
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
High-resolution US of non-traumatic recurrent dislocation of the peroneal tendons: a case report
err1998-06-19
err0
PREAI
errG. M. Magnano; Mauro Occhi; Mauro Di Stadio; Paolo Toma'; Lorenzo E. Derchi
err分享
err收藏
Heterogeneous Excitation-and-Squeeze Network for visual dialog
err2021-08-01
err7
PREAI
errLin, Bingqian; Zhu, Yi; Liang, Xiaodan
err分享
err收藏
Aligning vision-language for graph inference in visual dialog
err2021-12-01
err7
PREAI
errJiang, Tianling; Shao, Hailin; Tian, Xin; Ji, Yi; Liu, Chunping
err分享
err收藏
Recurrent dislocation of the peroneal tendons
err1986-03-01
err0
PREAI
errMarc A. Martens; Jan F. Noyez; Jozef C. Mulier
err分享
err收藏
学者 查看更多内容