arrow
返回

Multi-Modal Validation and Domain Interaction Learning for Knowledge-Based Visual Question Answering

delete2024-11-01
delete0
PRE
AI
徐宁 封面图
徐宁 (Ning Xu)
Y
Yifei Gao
刘
刘安安 (An-An Liu) *
田宏硕 封面图
田宏硕 (Hongshuo Tian)
张
张勇东 (Yongdong Zhang)
DOI:10.1109/TKDE.2024.3384270delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Knowledge-based Visual Question Answering (KB-VQA) aims to answer the image-aware question via the external knowledge, which requires an agent to not only understand images but also explicitly retrieve and integrate knowledge facts. Intuitively, to accurately answer the question, we humans can validate the retrieved knowledge based on our memory, and then align the knowledge facts with the image regions to infer answers. However, most existing methods ignore the process of knowledge validation and alignment. In this paper, we propose the Multi-Modal Validation and Domain Interaction Learning method, which consists of two components: 1) Multi-modal validation for knowledge retrieval. We propose the multi-modal validation module (MMV) to evaluate the confidence of each retrieved knowledge fact via images and questions, which preserves knowledge candidates effective for inferring answers. 2) Domain interaction for knowledge integration. We propose the Domain Interaction TRansformer module (DI-TR) to align visual regions with knowledge facts by the interaction learning in the improved transformer. Specifically, the inter-domain and intra-domain masks are injected into each self-attention layer to control the integration scope. The proposed method outperforms several strong baselines on three widely-used knowledge-based datasets: KRVQA, OK-VQA and VQA2.0. Extensive experiments and ablation studies demonstrate the effectiveness of multi-modal knowledge validation and domain interaction learning.
Keyword:
Visualization
Knowledge based systems
Transformers
Databases
Knowledge graphs
Question answering (information retrieval)
Task analysis
Multi-modal validation
domain interaction learning
knowledge-based visual question answering

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

T
tianjin university
学者数:
8.0W
论文数: 5.8W
被引数: 88
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

err分享
err收藏
Canadian physical activity guidelines for adults: are Canadians aware?
err2016-09-01
err0
errOAAI
errLeila Pfaeffli Dale; Allana G. LeBlanc; Krystn Orr; Tanya Berry; Sameer Deshpande; Amy E. Latimer-Cheung; Norm O’Reilly; Ryan E. Rhodes; Mark S. Tremblay; Guy Faulkner
err分享
err收藏
Cubic–Quartic Optical Soliton Perturbation with Differential Group Delay for the Lakshmanan–Porsezian–Daniel Model by Lie Symmetry
err2022-01-24
err0
errOAAI
errSachin Kumar; Anjan Biswas; Yakup Yıldırım; Luminita Moraru; Simona Moldovanu; Hashim M. Alshehri; Dalal Adnan Maturi; Dalal H. Al-Bogami
err分享
err收藏
Impact of Materials, Proportioning, and Curing on Ultra-High-Performance Concrete Properties
err2020-01-01
err0
PREAI
errAshley S. Carey; Isaac L. Howard; Dylan A. Scott; Robert D. Moser; Jay Shannon; Alta Knizley
err分享
err收藏
Long-Term Longitudinal Patterns of Patient-Reported Fatigue After Breast Cancer: A Group-Based Trajectory Analysis
err2022-07-01
err0
errOAAI
errInes Vaz-Luis; Antonio Di Meglio; Julie Havas; Mayssam El-Mouhebb; Pietro Lapidari; Daniele Presti; Davide Soldato; Barbara Pistilli; Agnes Dumas; Gwenn Menvielle; Cecile Charles; Sibille Everhard; Anne-Laure Martin; Paul H. Cottu; Florence Lerebours; Charles Coutant; Sarah Dauchy; Suzette Delaloge; Nancy U. Lin; Patricia A. Ganz; Ann H. Partridge; Fabrice André; Stefan Michiels
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
GRLC: Graph Representation Learning With Constraints
err2024-06-01
err27
PREAI
errPeng, Liang; Mo, Yujie; Xu, Jie; Shen, Jialie; Shi, Xiaoshuang; Li, Xiaoxiao; Shen, Heng Tao; Zhu, Xiaofeng
err分享
err收藏
学者 查看更多内容