arrow
返回

Inner Knowledge-based Img2Doc Scheme for Visual Question Answering

delete2022-03-04
delete12
PRE
AI
Q
Qun Li
F
Fu Xiao *
B
Bir Bhanu
B
Biyun Sheng
R
Richang Hong
DOI:10.1145/3489142delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual Question Answering (VQA) is a research topic of significant interest at the intersection of computer vision and natural language understanding. Recent research indicates that attributes and knowledge can effectively improve performance for both image captioning and VQA. In this article, an inner knowledge-based Img2Doc algorithm for VQA is presented. The inner knowledge is characterized as the inner attribute relationship in visual images. In addition to using an attribute network for inner knowledge-based image representation, VQA scheme is associated with a question-guided Doc2Vec method for question-answering. The attribute network generates inner knowledge-based features for visual images, while a novel question-guided Doc2Vec method aims at converting natural language text to vector features. After the vector features are extracted, they are combined with visual image features into a classifier to provide an answer. Based on our model, the VQA problem is resolved by textual question answering. The experimental results demonstrate that the proposed method achieves superior performance on multiple benchmark datasets.
Keyword:
VQA
dense image captioning
Doc2Vec
inner knowledge-based
attribute network

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

U
university of california riverside
学者数:
1.1W
论文数: 8.3K
被引数: 16
University of California System 封面图
University of California System
学者数:
37.7W
论文数: 33.8W
被引数: 6.6K
引用论文

引用论文

err2003-01-01
err0
PREAI
errTheo H.M. Smits; Bernard Witholt; Jan B. van Beilen
err分享
err收藏
err分享
err收藏
学者 查看更多内容