Return
Smooth diffusion model for multimodal recommendation
DOI:10.1016/j.knosys.2025.114807.png)
Abstract
En 中文
The multimodal recommendation integrates various item features, including visual, textual, and acoustic information, to enhance recommendation performance. However, multimodal information is often contaminated with noise. Recent research has primarily focused on the model perspective, actively injecting random noise and subsequently enhancing the model’s inherent denoising capability by capturing noise patterns. Nevertheless, multimodal features are usually high-dimensional, and the noise in high-dimensional data has complex structures, making it challenging to effectively capture noise patterns. In addition, due to the different dimensionalities of feature spaces across modalities, the impact of noise injected into these spaces varies correspondingly. To address these issues, we propose the smooth diffusion model for multimodal recommendation. Specifically, we combine the masked prediction paradigm with a diffusion model, randomly masking different parts of the data to help the model learn more noise patterns. We introduce a smoothing mechanism to align the effects of noise across different dimensions in the feature space, thereby preventing feature bias during the diffusion process. Additionally, we use user behavior patterns as guidance during inference to provide reasonable denoising. To achieve semantic alignment of multimodal information, we incorporate a contrastive learning mechanism, which can effectively denoise multimodal information. We conducted extensive experiments on three multimodal datasets to demonstrate the effectiveness of our method. Our code is available at github/SDMMR .
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

