arrow
Return

CoMAG: Collaborative GNN–LLM learning for multimodal attributed graphs

delete2026-08-18
delete0
PRE
AI
S
Sicheng Liang
Y
Yihao Wen
C
Chunpu Huang
Y
Yu Liu
Y
Yichen Nie
J
Jingqi Feng
Y
Yukai Huang
J
Jiawei Ye
吴杰 (Jie Wu) *
DOI:10.1016/j.knosys.2026.116864delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• CoMAG enables zero-shot reasoning on multimodal attributed graphs. • CoMAG combines graph-enhanced multimodal learning with frozen-MLLM reasoning. • Continuous graph embeddings and symbolic cues form a dual interface. • Label-free fusion and neighborhood smoothing improve prediction reliability. • CoMAG shows robust performance across four multimodal graph benchmarks. Abstract Multimodal Attributed Graphs (MAGs), which combine textual, visual, and structural information, are prevalent in domains such as e-commerce and social media. However, zero-shot reasoning on MAGs remains challenging. Existing graph neural networks can produce biased representations under modality imbalance and typically depend on task-specific supervision, whereas multimodal large language models (MLLMs) exhibit strong cross-modal reasoning ability but lack explicit modeling of graph structure. To address these challenges, we propose CoMAG, a collaborative framework that unifies graph-enhanced multimodal representation learning with frozen-MLLM reasoning. CoMAG extracts modality-specific embeddings via frozen encoders and refines them through self-supervised graph learning. It further derives discrete structural priors as symbolic tokens and aligns continuous graph representations with the frozen MLLM token space, enabling prompt-based inference without target labels or task-specific fine-tuning. A structure-aware reasoning layer then introduces salience-aware structural cues, label-free adaptive fusion, and neighborhood-consistent smoothing to improve reliability. Extensive experiments on real-world MAG datasets and multiple frozen MLLM backbones demonstrate that CoMAG achieves strong zero-shot performance across diverse graph domains. Further analyses confirm its robustness to modality perturbations, the complementary contributions of its graph–language interfaces, the transferability of its learned relational representations, and the limited overhead of its graph-side interface.
Keywords:
Multimodal attributed graphs
Graph neural networks
Large language models
Zero-shot learning
Multimodal reasoning

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

W
Waseda University
Scholars:
1.0W
Papers: 8.7K
Citations: 8.3K
F
fudan university
Scholars:
11.6W
Papers: 7.7W
Citations: 121