Return
Multi-modal Dual Attention Graph Contrastive Learning for Recommendation
DOI:10.1016/j.knosys.2026.115404.png)
Abstract
En 中文
Multi-modal recommender systems, incorporating rich content information (e.g., images and texts) into user behavior modeling, have attracted significant attention recently. Current work has successfully combined graph neural networks (GNNs) and contrastive learning to improve recommendation accuracy and mitigate the inherent sparse data problem. Yet, view augmentation strategies borrowed from other domains—such as edge or node dropout—tend to distort the original graph structure, leading to unintended semantic drift and suboptimal representation learning. Moreover, prior work has predominantly focused on optimizing inter-modal weights while overlooking user-specific modality preferences and adaptation of modal features generated by generic models. To tackle the above issues, we propose a novel multi-mOdal dUal aTtention Graph cOntrastive learning framework (OUTGO). Specifically, we first encode user and item representations by utilizing user and item homogeneous GNNs. Then, we employ designed intra- and inter-attention mechanisms, sequentially and adaptively, tuning each modal feature value based on the principal loss and considering fusing them with different modal perspectives. Additionally, semantic and structural contrastive learning tasks are introduced to alleviate the sparse data without destroying the original data structure. Extensive experiments on real-world datasets demonstrate the superiority of OUTGO compared to state-of-the-art baselines. The code is available at https://anonymous.4open.science/r/OUTGO .
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

