arrow
返回

Exploring Sparse Spatial Relation in Graph Inference for Text-Based VQA

delete2023-01-01
delete12
delete
OA
AI
S
Sheng Zhou
D
Dan Guo
J
Jia Li
X
Xun Yang *
王
王萌 (Meng Wang) *
DOI:10.1109/TIP.2023.3310332delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text-based visual question answering (TextVQA) faces the significant challenge of avoiding redundant relational inference. To be specific, a large number of detected objects and optical character recognition (OCR) tokens result in rich visual relationships. Existing works take all visual relationships into account for answer prediction. However, there are three observations: (1) a single subject in the images can be easily detected as multiple objects with distinct bounding boxes (considered repetitive objects). The associations between these repetitive objects are superfluous for answer reasoning; (2) two spatially distant OCR tokens detected in the image frequently have weak semantic dependencies for answer reasoning; and (3) the co-existence of nearby objects and tokens may be indicative of important visual cues for predicting answers. Rather than utilizing all of them for answer prediction, we make an effort to identify the most important connections or eliminate redundant ones. We propose a sparse spatial graph network (SSGN) that introduces a spatially aware relation pruning technique to this task. As spatial factors for relation measurement, we employ spatial distance, geometric dimension, overlap area, and DIoU for spatially aware pruning. We consider three visual relationships for graph learning: object-object, OCR-OCR tokens, and object-OCR token relationships. SSGN is a progressive graph learning architecture that verifies the pivotal relations in the correlated object-token sparse graph, and then in the respective object-based sparse graph and token-based sparse graph. Experiment results on TextVQA and ST-VQA datasets demonstrate that SSGN achieves promising performances. And some visualization results further demonstrate the interpretability of our method.
Keyword:
Visual question answering
text-based visual question answering
graph inference
spatial relation
relation learning

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

H
hefei university of technology
学者数:
2.5W
论文数: 1.7W
被引数: 35
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

A RARE INVOLVEMENT OF CENTRAL NERVOUS SYSTEM INVOLVEMENT DUE TO CTLA-4 GENE DEFECT
err2021-01-01
err0
errOAAI
errSahib Rovshanov; Rahşan Göçmen; Ibrahim Barışta; Deniz Çağdaş Ayvaz; Ayşegül Üner; Vedat Çilingir; İrsel Tezer; Çağman Tan; Elif Soyak Aytekin; İlhan Tezcan; Pinar Acar Özen; Aslı Tuncer
err分享
err收藏
Measuring static thermal permeability and inertial factor of rigid porous materials (L)
err2011-11-16
err0
PREAI
errM. Sadouki; M. Fellah; Z. E. A. Fellah; E. Ogam; N. Sebaa; F. G. Mitri; C. Depollier
err分享
err收藏
Risk of fingolimod rebound after switching to cladribine or rituximab in multiple sclerosis
err2022-06-01
err0
PREAI
errGro Owren Nygaard; Hilde Torgauten; Lars Skattebøl; Einar August Høgestøl; Piotr Sowa; Kjell-Morten Myhr; Øivind Torkildsen; Elisabeth Gulowsen Celius
err分享
err收藏
err分享
err收藏
学者 查看更多内容