arrow
Return

MESA: A Multimodal Entity Entailment framework for multimodal Entity Alignment

delete2025-01-01
delete0
PRE
AI
Y
Yu Zhao
Y
Ying Zhang *
X
Xuhui Sui
X
Xiangrui Cai
DOI:10.1016/j.ipm.2024.103951delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Current methods for multimodal entity alignment (MEA) primarily rely on entity representation learning, which undermines entity alignment performance because of cross-KG interaction deficiency and multimodal heterogeneity. In this paper, we propose a M ultimodal E ntity E ntailment framework of multimodal E ntity A lignment task, ME3A, and recast the MEA task as an entailment problem about entities in the two KGs. This way, the cross-KG modality information directly interacts with each other in the unified textual space. Specifically, we construct the multimodal information in the unified textual space as textual sequences: for relational and attribute modalities, we combine the neighbors and attribute values of entities as sentences; for visual modality, we map the entity image as trainable prefixes and insert them into sequences. Then, we input the concatenated sequences of two entities into the pre-trained language model (PLM) as an entailment reasoner to capture the unified fine-grained correlation pattern of the multimodal tokens between entities. Two types of entity aligners are proposed to model the bi-directional entailment probability as the entity similarity. Extensive experiments conducted on nine MEA datasets with various modality combination settings demonstrate that our MESA effectively incorporates multimodal information and surpasses the performance of the state-of-the-art MEA methods by 16.5% at most.
Keywords:
Entity Alignment
Multimodal learning
Knowledge graph
Prompt learning

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

N
nankai university
Scholars:
4.7W
Papers: 3.2W
Citations: 74