返回
Multi-aware coreference relation network for visual dialog
DOI:10.1007/s13735-022-00257-2.png)
摘要
En 中文
As a challenging cross-media task, visual dialog assesses whether an AI agent can converse in human language based on its understanding of visual content. So the critical issue is to pay attention not only to the problem of coreference in vision, but also to the problem of coreference in and between vision and language. In this paper, we propose the multi-aware coreference relation network (MACR-Net) to solve it from both textual and visual perspectives and to do fusion in complementary awareness. Specifically, its textual coreference relation module identifies textual coreference relations based on multi-aware textual representation from textual view. Furthermore, the visual coreference relation module adaptively adjusts visual coreference relations based on contextual-aware relations representation from visual view. Finally, the multi-modals fusion module fuses multi-aware relations to get an aligned representation. Extensive experiments on the VisDial v1.0 benchmarks show that MACR-Net achieves state-of-the-art performance.
Keyword:
Visual dialog
Multimedia
Coreference resolution
Cross-modal relationships
期刊
IF:
2.9
论文数:
278
被引数:
866
机构
引用论文
Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue视觉对话中自适应视觉理解的学习双编码模型
Cobalt and Copper Composite Oxides as Efficient Catalysts for Preferential Oxidation of CO in H2-Rich Stream钴和铜复合氧化物作为H2-Rich流中CO优先氧化的有效催化剂

