arrow
Return

Visual Spatial Reasoning

delete2023-06-20
delete25
delete
OA
AI
F
Fangyu Liu *
G
Guy Emerson
N
Nigel Collier
DOI:10.1162/tacl_a_00566delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Spatial relations are a basic part of human cognition. However, they are expressed in natural language in a variety of ways, and previous work has suggested that current vision-and-language models (VLMs) struggle to capture relational information. In this paper, we present Visual Spatial Reasoning (VSR), a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing). While using a seemingly simple annotation format, we show how the dataset includes challenging linguistic phenomena, such as varying reference frames. We demonstrate a large gap between human and model performance: The human ceiling is above 95%, while state-of-the-art models only achieve around 70%. We observe that VLMs' by-relation performances have little correlation with the number of training examples and the tested models are in general incapable of recognising relations concerning the orientations of objects.(1)

Journal

T
Transactions of the Association for Computational Linguistics
IF:
6.9
Papers:
486
Citations:
5.7K

Organization

U
University of Cambridge
Scholars:
7.7W
Papers: 7.1W
Citations: 13.7W