Return
Question-guided multigranular visual augmentation for knowledge-based visual question answering
DOI:10.1016/j.cviu.2025.104569.png)
Abstract
En 中文
• We design a novel question-guided multigranular visual feature extraction method to explore full and thorough interactions between input question and image. • Our method serves questions as guidance to locate the visual regions of interest for QA. • In our method, the convolutional kernels are dynamically derived from the question, which is effective in model parameters. • Our method achieves top performance over 2 knowledge-based VQA benchmarks.
Journal
IF:
3.5
Papers:
428
Citations:
7.3K

