arrow
Return

Knowledge-Embedded Mutual Guidance for Visual Reasoning

delete2024-04-01
delete1
PRE
AI
W
Wenbo Zheng *
L
Lan Yan
L
Long Chen
Q
Qiang Li
F
Fei‐Yue Wang
DOI:10.1109/TCYB.2023.3310892delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual reasoning between visual images and natural language is a long-standing challenge in computer vision. Most of the methods aim to look for answers to questions only on the basis of the analysis of the offered questions and images. Other approaches treat knowledge graphs as flattened tables to search for the answer. However, there are two major problems with these works: 1) the model disregards the fact that the world we surrounding us interlinks our hearing and speaking of natural language and 2) the model largely ignores the structure of the KG. To overcome these challenging deficiencies, a model should jointly consider two modalities of vision and language, as well as the rich structural and logical information embedded in knowledge graphs. To this end, we propose a general joint representation learning framework for visual reasoning, namely, knowledge-embedded mutual guidance. It realizes mutual guidance not only between visual data and natural language descriptions but also between knowledge graphs and reasoning models. In addition, it exploits the knowledge derived from the reasoning model to boost knowledge graphs when applying the visual relation detection task. The experimental results demonstrate that the proposed approach performs dramatically better than state-of-the-art methods on two benchmarks for visual reasoning.
Keywords:
Attention model
joint learning
knowledge embedding
visual reasoning

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

I
institute of automation, cas
Scholars:
2.2K
Papers: 2.1K
Citations: 2
H
hunan university
Scholars:
4.5W
Papers: 3.3W
Citations: 70
W
Wuhan University of Technology
Scholars:
3.4W
Papers: 2.4W
Citations: 4.4W
researcher View more organizations