arrow
返回

Knowledge is power: Open-world knowledge representation learning for knowledge-based visual reasoning ☆,☆☆ , ☆☆

delete2024-08-01
delete2
PRE
AI
W
Wenbo Zheng
L
Lan Yan
F
Fei‐Yue Wang
DOI:10.1016/j.artint.2024.104147delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Knowledge-based visual reasoning requires the ability to associate outside knowledge that is not present in a given image for cross-modal visual understanding. Two deficiencies of the existing approaches are that (1) they only employ or construct elementary and explicit but superficial knowledge graphs while lacking complex and implicit but indispensable cross-modal knowledge for visual reasoning, and (2) they also cannot reason new/ unseen images or questions in open environments and are often violated in real-world applications. How to represent and leverage tacit multimodal knowledge for open-world visual reasoning scenarios has been less studied. In this paper, we propose a novel open-world knowledge representation learning method to not only construct implicit knowledge representations from the given images and their questions but also enable knowledge transfer from a known given scene to an unknown scene for answer prediction. Extensive experiments conducted on six benchmarks demonstrate the superiority of our approach over other state-of-the-art methods. We apply our approach to other visual reasoning tasks, and the experimental results show that our approach, with its good performance, can support related reasoning applications.
Keyword:
Visual reasoning
Knowledge representation learning
Open-world learning
Graph model

期刊

Artificial Intelligence Review 封面图
Artificial Intelligence Review
IF:
13.9
论文数:
6.1K
被引数:
1.9W

机构

H
hunan university
学者数:
4.5W
论文数: 3.3W
被引数: 70
W
Wuhan University of Technology
学者数:
3.4W
论文数: 2.4W
被引数: 4.4W
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

Transformers in Vision: A Survey视觉中的变形金刚: 一项调查
err2022-09-13
err1.1K
errOAAI
errKhan, Salman; Naseer, Muzammal; Hayat, Munawar; Zamir, Syed Waqas; Khan, Fahad Shahbaz; Shah, Mubarak
err分享
err收藏
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Fact-based visual question answering via dual-process system
err2022-02-01
err17
PREAI
errLiu, Luping; Wang, Meiling; He, Xiaohai; Qing, Linbo; Chen, Honggang
err分享
err收藏
学者 查看更多内容