Return
Adaptive Graph Convolution With Diffusion Models for Multimodal Recommendation
J
Z
B
P
DOI:10.1109/tkde.2026.3706381.png)
Abstract
En 中文
Multimodal recommendation systems (MRSs) address data sparsity and cold-start issues in collaborative filtering by incorporating multimodal item content. Recent MRSs explicitly construct item–item graphs from modality feature similarity and employ graph convolutional networks (GCNs) to propagate information for learning item representations. However, item–item graphs constructed from a single source may fail to capture complementary relationships across different information sources, while fixed propagation rules tend to blur fine-grained semantic distinctions, resulting in suboptimal recommendations. To overcome these limitations, we propose DiffGCN, a novel framework that leverages conditional diffusion models (DMs) to enable adaptive graph convolution on item–item graphs. Specifically, DiffGCN learns item-specific weights to adaptively fuse multi-source neighborhoods. On this basis, DiffGCN replaces the one-step propagation of conventional GCNs with a multi-step denoising process conditioned on neighborhood information. This iterative refinement adaptively integrates multi-source neighborhood information while preserving fine-grained semantics, yielding more expressive item representations. Extensive experiments on three real-world datasets demonstrate that DiffGCN consistently outperforms baselines while maintaining scalability, validating the effectiveness of diffusion-aware adaptive graph convolution in multimodal recommendation.
Keywords:
Recommender system
multimodal recommendation
diffusion model
graph convolutional network
Journal
IF:
10.4
Papers:
6.7K
Citations:
3.2W
