1
Return

Evolution Rather Than Degradation: Structure-Guided Elastic Consensus Learning for Multimodal Knowledge Graph Completion

delete2026-07-06
delete0
PRE
AI
Y
Yameng Liu
S
Shuai Zheng
Z
Zhenfeng Zhu
Y
Yunhui Xu
Y
Yan Zhuang
赵耀 (Yao Zhao)
K
Kunlun He
DOI:10.1109/tkde.2026.3710321delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent efforts in multimodal knowledge graph completion (MKGC) have demonstrated impressive progress empowered by contrastive learning. However, a non-common sense observation shows that the simple pursuit of cross-modal consensus might degenerate MKGC performance unexpectedly, owing to the integration of untrustworthy auxiliary modalities. This behavior is herein referred to as modality degradation, which inadvertently diminishes the synergy between contrastive learning and MKGC. In this paper, we introduce a <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">S</u>tructure-Guided <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E</u>lastic <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">C</u>onsensus <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">L</u>earning (SECL) framework. It captures cross-modal consensus from trustworthy multimodal entity representations for MKGC. Specifically, to reduce the semantic bias observed in CLIP-derived multimodal representations, a structure-guided multimodal representation harmonization method is introduced in SECL. With the structural context of different entities being a coordinator, this method selectively suppresses irrelevant features in CLIP-derived multimodal representations, thereby promoting their inter-modal semantic consistency. To supervise this harmonization process, an elastic contrastive learning method is presented. It alleviates the under-exploration of cross-modal consensus by steering the model optimization to prioritize aligning multimodal representations for hard entities. Extensive experiments on benchmark datasets show that SECL exhibits modality evolution instead of degradation and yields substantial improvements over state-of-the-art methods.
Keywords:
Multimodal knowledge graph
knowledge graph completion
contrastive learning
modality alignment

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

C
chinese pla general hospital
Scholars:
1.5K
Papers: 475
Citations: 4
B
Beijing Jiaotong University
Scholars:
2.1W
Papers: 1.7W
Citations: 1.2W
Cited Papers

Cited Papers

Citing Papers

Citing Papers