返回
Knowledge is power: Open-world knowledge representation learning for knowledge-based visual reasoning ☆,☆☆ , ☆☆
DOI:10.1016/j.artint.2024.104147.png)
摘要
En 中文
Knowledge-based visual reasoning requires the ability to associate outside knowledge that is not present in a given image for cross-modal visual understanding. Two deficiencies of the existing approaches are that (1) they only employ or construct elementary and explicit but superficial knowledge graphs while lacking complex and implicit but indispensable cross-modal knowledge for visual reasoning, and (2) they also cannot reason new/ unseen images or questions in open environments and are often violated in real-world applications. How to represent and leverage tacit multimodal knowledge for open-world visual reasoning scenarios has been less studied. In this paper, we propose a novel open-world knowledge representation learning method to not only construct implicit knowledge representations from the given images and their questions but also enable knowledge transfer from a known given scene to an unknown scene for answer prediction. Extensive experiments conducted on six benchmarks demonstrate the superiority of our approach over other state-of-the-art methods. We apply our approach to other visual reasoning tasks, and the experimental results show that our approach, with its good performance, can support related reasoning applications.
Keyword:
Visual reasoning
Knowledge representation learning
Open-world learning
Graph model
期刊
IF:
13.9
论文数:
6.1K
被引数:
1.9W
机构
引用论文
Proceedings of the 4th International Conference on the Industry 4.0 Model for Advanced Manufacturing
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
Image captioning for effective use of language models in knowledge-based visual question answering在基于知识的视觉问答中有效使用语言模型的图像字幕

