Return
Missing visual modality graph transformer for multi-modal entity alignment
DOI:10.1016/j.patcog.2026.113320.png)
Abstract
En 中文
• Propose MVMTEA, a novel approach to deal with the problem of missing visual modality in multimodal entity alignment. • A dual-level modality completion module is proposed for missing visual modality, including Dirichlet energy minimisation-based neighborhood modality propagation (local-level completion) and KNN-based correlation aggregation (global-level completion). • MVMTEA introduces transformer fine-grain to achieve the aggregation of different modalities and contrastive learning to achieve entity-level alignment. • Extensive experiments on three bilingual DBP15K datasets are conducted to demonstrate that MVMTEA is suitable for visual modality missing scenarios and outperforms eleven baseline methods.
Keywords:
MVMTEA
multimodal entity alignment
missing visual modality
graph transformer
contrastive learning

