arrow
返回

Self-Supervised Learning for Multimedia Recommendation

delete2023-01-01
delete28
PRE
AI
Z
Zhulin Tao
L
Liu, XH
王翔 (Xiang Wang) *
L
Lifang Yang
X
Xianglin Huang
T
Tat‐Seng Chua
DOI:10.1109/TMM.2022.3187556delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Learning representations for multimedia content is critical for multimedia recommendation. Current representation learning methods roughly fall into two groups: (1) using the historical interactions to create ID embeddings of users and items, and (2) treating multi-modal data as the side information of items to enrich their ID embeddings. Each user-item interaction offers the supervisory signal to optimize the representation learning by the traditional supervised learning paradigm. Due to the overlook of the multi-modal patterns ($e.g.$, co-occurrence of visual, acoustic, textual features in micro-videos a user saw before, and her behavioral features) hidden in the data, these methods are insufficient to create powerful representations and obtain satisfactory recommendation accuracy. To capture multi-modal patterns in the data itself, we go beyond the supervised learning paradigm, and incorporate the idea of self-supervised learning (SSL) into multimedia recommendation. Specifically, SSL consists of two components: (1) data augmentation upon multi-modal contents, where we design three operators - feature dropout (FD), feature masking (FM), feature fine and coarse spaces (FAC) - to generate multiple views of individual items; and (2) contrastive learning, which differentiates the views of an item from the others' to distill additional supervisory signals. Clearly, SSL enables us to explore and exhibit the underlying relations among modalities, thereby resulting in powerful representations. We denote the generic framework by Self-supervised Learning-guided Multimedia Recommendation (SLMRec). Extensive experiments are performed on three real-world datasets, showing that SLMRec achieves significant improvements over several state-of-the-art baselines like LightGCN [1], MMGCN [2]. Further analysis shows how SSL affects recommendation performance.
Keyword:
Multimedia recommendation
self-supervised learning
graph neural network
micro-videos

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

C
Communication University of China
学者数:
1.1K
论文数: 820
被引数: 326
U
university of science & technology of china, cas
学者数:
3.2W
论文数: 2.7W
被引数: 74
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
HoAFM: A High-order Attentive Factorization Machine for CTR Prediction
err2020-11-01
err49
PREAI
errTao, Zhulin; Wang, Xiang; He, Xiangnan; Huang, Xianglin; Chua, Tat-Seng
err分享
err收藏
Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image Captioning
err2020-05-01
err78
PREAI
errXu, Ning; Zhang, Hanwang; Liu, An-An; Nie, Weizhi; Su, Yuting; Nie, Jie; Zhang, Yongdong
err分享
err收藏
Hierarchical User Intent Graph Network for Multimedia Recommendation
err2022-01-01
err42
PREAI
errWei, Yinwei; Wang, Xiang; He, Xiangnan; Nie, Liqiang; Rui, Yong; Chua, Tat-Seng
err分享
err收藏
Knowledge-Based Topic Model for Multi-Modal Social Event Analysis
err2020-08-01
err23
PREAI
errXue, Feng; Hong, Richang; He, Xiangnan; Wang, Jianwei; Qian, Shengsheng; Xu, Changsheng
err分享
err收藏
err分享
err收藏
err分享
err收藏
Individual and combining effects of anti-RANKL monoclonal antibody and teriparatide in ovariectomized mice
err2015-06-01
err0
errOAAI
errNaoto Tokuyama; Jun Hirose; Yasunori Omata; Tetsuro Yasui; Naohiro Izawa; Takumi Matsumoto; Hironari Masuda; Toshinobu Ohmiya; Hisataka Yasuda; Taku Saito; Yuho Kadono; Sakae Tanaka
err分享
err收藏
Highly Polymorphic G-quadruplexes in the c-MYC Promoter
err2010-04-20
err0
errOAAI
errJeong-Min Yoon; Hyun-Jin Kang; Jae-Ho Sung; Hyun-Ju Park; Sung-Chul Hohng
err分享
err收藏
学者 查看更多内容