arrow
返回

Transformer-Based Relational Inference Network for Complex Visual Relational Reasoning

delete2023-08-25
delete1
PRE
AI
谭明奎 封面图
谭明奎 (Mingkui Tan) *
Z
Zhiquan Wen
方乐缘 封面图
方乐缘 (Leyuan Fang)
Q
Qi Wu
DOI:10.1145/3605781delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual Relational Reasoning is the basis of many vision-and-language based tasks (e.g., visual question answering and referring expression comprehension). In this article, we regard the complex referring expression comprehension (c-REF) task as the reasoning basis, in which c-REF seeks to localise a target object in an image guided by a complex query. Such queries often contain complex logic and thus impose two critical challenges for reasoning: (i) Comprehending the complex queries is difficult since these queries usually refer to multiple objects and their relationships; (ii) Reasoning among multiple objects guided by the queries and then localising the target correctly are non-trivial. To address the above challenges, we propose a Transformer-based Relational Inference Network (Trans-RINet). Specifically, to comprehend the queries, we mimic the language-comprehending mechanism of humans, and devise a language decomposition module to decompose the queries into four types, i.e., basic attributes, absolute location, visual relationship and relative location. We further devise four modules to address the corresponding information. In each module, we consider the intra-(i.e., between the objects) and inter-modality relationships(i.e., between the queries and objects) to improve the reasoning ability. Moreover, we construct a relational graph to represent the objects and their relationships, and devise a multi-step reasoning method to progressively understand the complex logic. Since each type of the queries is closely related, we let each module interact with each other before making a decision. Extensive experiments on the CLEVR-Ref+, Ref-Reasoning, and CLEVR-CoGenT datasets demonstrate the superior reasoning performance of our Trans-RINet.
Keyword:
Visual Relational Reasoning
complex referring expression comprehension
Gated Graph Neural Network

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

U
University of Adelaide
学者数:
2.3W
论文数: 2.4W
被引数: 4.2W
H
hunan university
学者数:
4.5W
论文数: 3.3W
被引数: 70
S
south china university of technology
学者数:
6.8W
论文数: 5.1W
被引数: 85
学者 查看更多机构
引用论文

引用论文

A high temperature variety of BiOF一种高温品种的BiOF
err1983-09-01
err0
PREAI
errSamir Matar; Jean-Maurice Reau; Louis Rabardel; Gérard Demazeau; Paul Hagenmuller
err分享
err收藏
没有更多内容