arrow
返回

Visual Grounding With Joint Multimodal Representation and Interaction

delete2023-01-01
delete0
PRE
AI
H
Hong Zhu
Q
Qingyang Lu
L
Lei Xue
X
Xue, Mogen
G
Guanglin Yuan *
B
Bineng Zhong
DOI:10.1109/TIM.2023.3324362delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This article tackles the challenging yet significant task of grounding a natural language query to the corresponding region onto an image. The main challenge in visual grounding is to model the correspondence between visual context and semantic concept referred by the language expression, i.e., multimodal fusion. Nevertheless, there is an inherent deficiency in the current fusion module designs, which makes visual and linguistic feature embeddings cannot be unified into the same semantic space. To address the issue, we present a novel and effective visual grounding framework based on joint multimodal representation and interaction (JMRI). Specifically, we propose to perform image-text alignment in a multimodal embedding space learned by a large-scale foundation model, so as to obtain semantically unified joint representations. Furthermore, the transformer-based deep interactor is designed to capture intramodal and intermodal correlations, rendering our model to highlight the localization-relevant cues for accurate reasoning. By freezing the pretrained vision-language foundation model and updating the other modules, we achieve the best performance with the lowest training cost. Extensive experimental results on five benchmark datasets with quantitative and qualitative analysis show that the proposed method performs favorably against the state-of-the-arts.
Keyword:
Cross-modal interaction
feature alignment
image-text foundation model
visual grounding

期刊

IEEE Transactions on Instrumentation and Measurement 封面图
IEEE Transactions on Instrumentation and Measurement
IF:
5.9
论文数:
1.9W
被引数:
5.8W

机构

G
Guangxi Normal University
学者数:
7.7K
论文数: 4.9K
被引数: 5.1K
N
national university of defense technology - china
学者数:
1.8W
论文数: 1.4W
被引数: 9
引用论文

引用论文

Enhanced mechanical behaviour of lead zirconate titanate piezoelectric composites incorporating zinc oxide nanowhiskers
err2008-11-25
err0
PREAI
errLin Hai-Bo; Cao Mao-Sheng; Yuan Jie; Wang Da-Wei; Zhao Quan-Liang; Wang Fu-Chi
err分享
err收藏
err分享
err收藏
Climate change discourse among Iranian farmers
err2016-07-16
err0
PREAI
errTahereh Zobeidi; Masoud Yazdanpanah; Masoumeh Forouzani; Bahman Khosravipour
err分享
err收藏
学者 查看更多内容