arrow
Return

Multimodal Graph Contrastive Learning for Multimedia-Based Recommendation

delete2023-01-01
delete9
PRE
AI
刘康 cover
刘康 (Kang Liu)
F
Feng Xue *
D
Dan Guo
P
Peijie Sun
钱胜胜 (Shengsheng Qian)
R
Richang Hong
DOI:10.1109/TMM.2023.3251108delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimedia-based recommendation is a challenging task that requires not only learning collaborative signals from user-item interaction, but also capturing modality-specific user interest clues from complex multimedia content. Though significant progress on this challenge has been made, we argue that current solutions remain limited by multimodal noise contamination. Specifically, a considerable proportion of multimedia content is irrelevant to the user preference, such as the background, overall layout, and brightness of images; the word order and semantic-free words in titles; etc. We take this irrelevant information as noise contamination to discover user preferences. Moreover, most recent research has been conducted by graph learning. This means that noise is diffused into the user and item representations with the message propagation; the contamination influence is further amplified. To tackle this problem, we develop a novel framework named Multimodal Graph Contrastive Learning (MGCL), which captures collaborative signals from interactions and uses visual and textual modalities to respectively extract modality-specific user preference clues. The key idea of MGCL involves two aspects: First, to alleviate noise contamination during graph learning, we construct three parallel graph convolution networks to independently generate three types of user and item representations, containing collaborative signals, visual preference clues, and textual preference clues. Second, to eliminate as much preference-independent noisy information as possible from the generated representations, we incorporate sufficient self-supervised signals into the model optimization with the help of contrastive learning, thus enhancing the expressiveness of the user and item representations. Extensive experiments validate the effectiveness and scalability of MGCL at https://github.com/hfutmars/MGCL.
Keywords:
Collaboration
Visualization
Contamination
Convolution
Neural networks
Feature extraction
Electronic mail
Recommender system
collaborative filtering
multimodal user preference
graph convolution network
contrastive learning

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

H
hefei university of technology
Scholars:
2.5W
Papers: 1.7W
Citations: 35
I
institute of automation, cas
Scholars:
2.2K
Papers: 2.1K
Citations: 2
C
chinese academy of sciences
Scholars:
56.3W
Papers: 44.8W
Citations: 704
researcher View more organizations