返回
Vision-Language-Knowledge Co-Embedding for Visual Commonsense Reasoning
DOI:10.3390/s21092911.png)
摘要
En 中文
Visual commonsense reasoning is an intelligent task performed to decide the most appropriate answer to a question while providing the rationale or reason for the answer when an image, a natural language question, and candidate responses are given. For effective visual commonsense reasoning, both the knowledge acquisition problem and the multimodal alignment problem need to be solved. Therefore, we propose a novel Vision-Language-Knowledge Co-embedding (ViLaKC) model that extracts knowledge graphs relevant to the question from an external knowledge base, ConceptNet, and uses them together with the input image to answer the question. The proposed model uses a pretrained vision-language-knowledge embedding module, which co-embeds multimodal data including images, natural language texts, and knowledge graphs into a single feature vector. To reflect the structural information of the knowledge graph, the proposed model uses the graph convolutional neural network layer to embed the knowledge graph first and then uses multi-head self-attention layers to co-embed it with the image and natural language question. The effectiveness and performance of the proposed model are experimentally validated using the VCR v1.0 benchmark dataset.
Keyword:
visual commonsense reasoning
multimodal co-embedding
knowledge graph
graph convolutional network
pretrained multi-head self-attention network
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.5
论文数:
7.2W
被引数:
20.9W
机构
引用论文
Kinematic analysis of limb movements in neuropsychological research: Subtle deficits and recovery of function.神经心理学研究中肢体运动的运动学分析: 细微的缺陷和功能的恢复。
Effect of Lizards on Spider Populations: Manipulative Reconstruction of a Natural Experiment
Science
IF0
Two-dimensional bricklayer arrangements of tolans using halogen bonding interactions使用卤素键相互作用的tolans的二维瓦工层布置

