arrow
Return

Question-guided multigranular visual augmentation for knowledge-based visual question answering

delete2025-11-20
delete0
PRE
AI
刘景 (Jing Liu)
L
Lizong Zhang *
C
Chong Mu
L
LU Guang-xi
B
Ben Zhang
J
Junsong Li
DOI:10.1016/j.cviu.2025.104569delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We design a novel question-guided multigranular visual feature extraction method to explore full and thorough interactions between input question and image. • Our method serves questions as guidance to locate the visual regions of interest for QA. • In our method, the convolutional kernels are dynamically derived from the question, which is effective in model parameters. • Our method achieves top performance over 2 knowledge-based VQA benchmarks.

Journal

Computer Vision and Image Understanding cover
Computer Vision and Image Understanding
IF:
3.5
Papers:
428
Citations:
7.3K

Organization

U
university of electronic science and technology of china
Scholars:
1.2W
Papers: 4.5K
Citations: 4