返回
Visual Grounding With Joint Multimodal Representation and Interaction
DOI:10.1109/TIM.2023.3324362.png)
摘要
En 中文
This article tackles the challenging yet significant task of grounding a natural language query to the corresponding region onto an image. The main challenge in visual grounding is to model the correspondence between visual context and semantic concept referred by the language expression, i.e., multimodal fusion. Nevertheless, there is an inherent deficiency in the current fusion module designs, which makes visual and linguistic feature embeddings cannot be unified into the same semantic space. To address the issue, we present a novel and effective visual grounding framework based on joint multimodal representation and interaction (JMRI). Specifically, we propose to perform image-text alignment in a multimodal embedding space learned by a large-scale foundation model, so as to obtain semantically unified joint representations. Furthermore, the transformer-based deep interactor is designed to capture intramodal and intermodal correlations, rendering our model to highlight the localization-relevant cues for accurate reasoning. By freezing the pretrained vision-language foundation model and updating the other modules, we achieve the best performance with the lowest training cost. Extensive experimental results on five benchmark datasets with quantitative and qualitative analysis show that the proposed method performs favorably against the state-of-the-arts.
Keyword:
Cross-modal interaction
feature alignment
image-text foundation model
visual grounding
期刊
IF:
5.9
论文数:
1.9W
被引数:
5.8W
机构
引用论文
Substrate specificity and inhibitor analyses of human steroid 5β-reductase (AKR1D1)人类类固醇5 β-还原酶 (AKR1D1) 的底物特异性和抑制剂分析
Steroids
IF0
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?
RI-Fusion: 3D Object Detection Using Enhanced Point Features With Range-Image Fusion for Autonomous DrivingRI融合: 使用增强的点特征进行3D物体检测,并进行距离图像融合,以实现自动驾驶
Communication and reproductive behaviour in North American Jerusalem crickets (
Stenopelmatus
) (Orthoptera: Stenopelmatidae).北美耶路撒冷蟋蟀的交流和生殖行为 (
Stenopelmatus
) (直翅目: Stenopelmatidae)。
Generation of hydroxyl radicals by urban suspended particulate air matter. The role of iron ions城市悬浮颗粒物产生羟基自由基。铁离子的作用:

