arrow
返回

Graph Pooling Inference Network for Text-based VQA

delete2024-01-11
delete7
PRE
AI
S
Sheng Zhou
D
Dan Guo
X
Xun Yang *
J
Jianfeng Dong
王萌 (Meng Wang) *
DOI:10.1145/3634918delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Effectively leveraging objects and optical character recognition (OCR) tokens to reason out pivotal scene text is critical for the challenging Text-based Visual Question Answering (TextVQA) task. Graph-based models can effectively capture the semantic relationship among visual entities (objects and tokens) and report remarkable performance in TextVQA. However, previous efforts usually leverage all visual entities and ignore the negative effect of superfluous entities. This article presents a Graph Pooling Inference Network (GPIN), which is an evolutionary graph learning method to purify the visual entities and capture the core semantics. It is observed that the dense distribution of reduplicative objects and the crowd of semantically dependent OCR tokens usually co-exist in the image. Motivated by this, GPIN adopts an adaptive node dropping strategy to dynamically downscale semantically closed nodes for graph evolution and update. To deepen the comprehension of scene text, GPIN is a dual-path hierarchical graph architecture that progressively aggregates the evolved object graph and the evolved token graph semantics into a graph vector that serves as visual cues to facilitate the answer reasoning. It can effectively eliminate object redundancy and enhance the association of semantically continuous tokens. Experiments conducted on TextVQA and ST-VQA datasets show that GPIN achieves promising performance compared with state-of-the-art methods.
Keyword:
Text-based visual question answering
graph inference
graph pooling

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

H
hefei university of technology
学者数:
2.5W
论文数: 1.7W
被引数: 35
U
university of science & technology of china, cas
学者数:
3.2W
论文数: 2.7W
被引数: 74
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
Video Moment Retrieval With Cross-Modal Neural Architecture Search
err2022-01-01
err69
PREAI
errYang, Xun; Wang, Shanshan; Dong, Jian; Dong, Jianfeng; Wang, Meng; Chua, Tat-Seng
err分享
err收藏
Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering
err2023-01-01
err6
errOAAI
errLi, Hao; Huang, Jinfa; Jin, Peng; Song, Guoli; Wu, Qi; Chen, Jie
err分享
err收藏
没有更多内容