arrow
返回

Fused modality-enhanced graph convolutional network for multimodal recommendation

delete2026-01-07
delete0
PRE
AI
徐
徐昊 (Hao Xu)
H
Hongbin Xia *
X
Xiaofeng Wang
DOI:10.1007/s10115-025-02652-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Data sparsity remains a critical challenge that hinders the performance of recommendation systems. Multimodal recommendation, which aims to alleviate this issue by enriching item and user representations through heterogeneous modality information, has garnered widespread attention. While existing studies have achieved notable success, their early or late fusion strategies inadequately balance modality-specific characteristics with cross-modal correlations. Moreover, static fusion paradigms are unable to effectively capture the dynamic changes in modality complementarity. To address these limitations, we propose a fused modality-enhanced graph convolutional network for multimodal recommendation (FM-GCN). Our method innovatively treats fused multimodal features as a distinct modality, achieving refined feature integration and noise suppression via a cross-attention mechanism and a self-prompted denoising diffusion model. Moreover, a decoupled feature distillation framework is employed to extract task-specific modality representations. Additionally, we construct a fused modality-enhanced user–item graph and a behavior-augmented item–item graph to synergistically model user interaction patterns and multimodal semantic relationships. Comprehensive experiments on four real-world datasets demonstrate that FM-GCN outperforms baselines with average improvements of 6.18%–8.10% in Recall@20 and NDCG@20 metrics, validating its effectiveness.
Keyword:
Multimodal recommendation
Feature fusion
Graph convolutional network
Knowledge distillation

期刊

Knowledge and Information Systems 封面图
Knowledge and Information Systems
IF:
3.1
论文数:
564
被引数:
5.2K

机构

S
School of Artificial Intelligence and Computer Science
学者数:
112
论文数: 43
被引数: 0
引用论文

引用论文

Recommender Systems Leveraging Multimedia Content
err2020-09-28
err111
PREAI
errDeldjoo, Yashar; Schedl, Markus; Cremonesi, Paolo; Pasi, Gabriella
err分享
err收藏
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
err2023-11-08
err1
PREAI
errLiu, Han; Wei, Yinwei; Liu, Fan; Wang, Wenjie; Nie, Liqiang; Chua, Tat-Seng
err分享
err收藏
Variable Kernel Density Estimation
err1992-09-01
err0
errOAAI
errGeorge R. Terrell; David W. Scott
err分享
err收藏
err分享
err收藏
Modeling Relational Data with Graph Convolutional Networks基于图卷积网络的关系数据建模
err2018-06-03
err0
PREAI
errMichael Schlichtkrull; Thomas N. Kipf; Peter Bloem; Rianne van den Berg; Ivan Titov; Max Welling
err分享
err收藏
学者 查看更多内容