arrow
Return

Visual question answering model based on multi-modal knowledge autonomous learning

delete2026-07-28
delete0
PRE
AI
Y
Yuanlong Wang *
Z
Zhiru Xu
H
Houshuai Wang
Z
Zhiwei Wu
H
Hu Zhang
DOI:10.1007/s11042-026-21802-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Knowledge-based visual question answering has made progress in cross-modal feature fusion and external knowledge integration, yet it still suffers from limitations in dynamic knowledge updating and continual learning, which constrain knowledge coverage and noise suppression. To address these issues, we propose a visual question answering model based on autonomous multi-modal knowledge learning, in which knowledge learning is divided into two stages: knowledge accumulation and knowledge updating. Specifically, large-scale multi-modal knowledge is first accumulated by constructing pseudo data, and then dynamically filtered and iteratively updated through an autonomous updating strategy to reduce noise and improve knowledge quality. During inference, the model actively retrieves and invokes relevant knowledge via a vector index, enabling automatic knowledge utilization and continual accumulation. Compared with the baseline models MuKEA and CMLR, the proposed method achieves accuracy improvements of 4.21% and 2.87% on the OK-VQA dataset, respectively, demonstrating that it effectively alleviates insufficient knowledge coverage and noise sensitivity in knowledge-based VQA and significantly enhances cross-modal understanding.
Keywords:
Knowledge-based visual question answering
Multi-modal knowledge
Autonomous learning
Vector indexing mechanism

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
1.9W
Citations:
3.2W

Organization

S
school of computer and information technology
Scholars:
17
Papers: 7
Citations: 0