arrow
返回

Multi-view semantic understanding for visual dialog

delete2023-05-01
delete1
PRE
AI
T
Tianling Jiang
Z
Zefan Zhang
X
Xin Li
季怡 封面图
季怡 (Yi Ji) *
刘纯平 封面图
刘纯平 (Chunping Liu)
DOI:10.1016/j.knosys.2023.110427delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual dialog, as a challenging cross-media task, requires answering a sequence of questions based on a given image and dialog history. Hence the key problem becomes how to answer visually grounded questions based on ambiguous reference information from dialog. In this work, we propose a novel method called Multi-View Semantic Understanding for Visual Dialog (MVSU) to resolve the visual coreference resolution problem. The model consists of two main textual processing modules, SRR (Semantic Retention RNN) and CRoT (Coreference Resolution on Text). Specifically, the SRR module generates word features that have semantical meaning by considering contextual information. The CRoT module is from a textual perspective to divide all useful nouns and pronouns into different clusters that serve as the supplement of the detailed information for semantic understanding. In experiments, we demonstrate that MVSU enhances the ability to understand the semantical information on the VisDial v1.0 dataset.
Keyword:
Visual dialog
Cross-media
Reference information
Semantic understanding

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

S
soochow university - china
学者数:
5.2W
论文数: 3.6W
被引数: 82
引用论文

引用论文

A novel intracochlear injection method for rapid drug delivery to vestibular end organs
err2020-07-01
err0
errOAAI
errVishal Raghu; Yugandhar Ramakrishna; Robert F. Burkard; Soroush G. Sadeghi
err分享
err收藏
NMN-VD: A Neural Module Network for Visual Dialog
errSENSORS
IF3.5
err2021-01-30
err4
errOAAI
errCho, Yeongsu; Kim, Incheol
err分享
err收藏
Reduced ACTH, while normal β-endorphin CSF levels in early epileptic encephalopathies
err1985-01-01
err0
PREAI
errF. Facchinetti; A. Nalin; F. Petraglia; V. Galli; A.R. Genazzani
err分享
err收藏
err分享
err收藏
学者 查看更多内容