arrow
返回

Encoder-decoder cycle for visual question answering based on perception-action cycle

delete2023-12-01
delete6
PRE
AI
A
Amin Jalali
M
Minho Lee *
DOI:10.1016/j.patcog.2023.109848delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this study, we propose a novel encoder-decoder cycle (EDC) framework inspired by the human learning process called the perception-action cycle to tackle challenging problems such as visual question answering (VQA) and visual relationship detection (VRD). EDC considers the understanding of the visual features of an image as perception and the act of answering the question regarding that image as an action. In the perception-action cycle, information is primarily collected from the environment and then passed to sensory structures in the brain to form an understanding of the environment. Acquired knowledge is then passed to motor structures to perform an action on the environment. Next, sensory structures perceive the altered environment and improve their understanding of the surrounding world. This process of understanding the environment, performing an action correspondingly, and then re-evaluating the initial understanding occurs cyclically in human life. EDC initially mimics this mechanism of introspection by comprehending and refining visual features to acquire the proper knowledge for answering the question. Subsequently, it decodes visual and language features into answer features, feeding them back cyclically to the encoder. In the VRD task, EDC decodes visual features to generate predicate features. We evaluate the proposed framework on the TDIUC, VQA 2.0, and VRD datasets, which outperforms the state-of-the-art models on the TDIUC and VRD datasets.
Keyword:
Visual question answering
Vision language tasks
Multi-modality fusion
Attention
Bilinear fusion
Brain-inspired frameworks

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

K
kyungpook national university (knu)
学者数:
1.8W
论文数: 1.8W
被引数: 14
引用论文

引用论文

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
err分享
err收藏
Generation of spin cat states in an engineered Dicke model
err2021-11-29
err0
errOAAI
errCaspar Groiseau; Stuart J. Masson; Scott Parkins
err分享
err收藏
Direct replication of Gervais & Norenzayan (2012): No evidence that analytic thinking decreases religious belief
err2017-02-24
err0
errOAAI
errClinton Sanchez; Brian Sundermeier; Kenneth Gray; Robert J. Calin-Jageman
err分享
err收藏
学者 查看更多内容