返回
Deep contrastive representation learning for multi-modal clustering
DOI:10.1016/j.neucom.2024.127523.png)
摘要
En 中文
Benefiting from the informative expression capability of contrastive representation learning (CRL), recent multi -modal learning studies have achieved promising clustering performance. However, it should be pointed out that the existing multi -modal clustering methods based on CRL fail to simultaneously take the similarity information embedded in inter- and intra-modal levels. In this study, we mainly explore deep multi -modal contrastive representation learning, and present a multi -modal learning network, named trustworthy multimodal contrastive clustering (TMCC), which incorporates contrastive learning and adaptively reliable sample selection with multi -modal clustering. Specifically, we are concerned with an adaptive filter to learn TMCC via progressing from 'easy' to 'complex' samples. Based on this, with the highly confident clustering labels, we present a new contrastive loss to learn modal -consensus representation, which takes into account not only the inter -modal similarity but also the intra-modal similarity. Experimental results show that these principles in TMCC consistently help promote clustering performance improvement.
Keyword:
Multi-view representation learning
Self-supervision
Clustering
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Synthesis of modified proanthocyanidins: easy and general introduction of a hydroxy group at C-6 of catechin; efficient synthesis of elephantorrhizol改性原花青素的合成: 在儿茶素C-6引入羟基; 象花素的高效合成
Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion ModelsDif融合: 在具有扩散模型的红外和可见光图像融合中实现高色彩保真度
Multimodal Weibull Variational Autoencoder for Jointly Modeling Image-Text Data用于联合建模图像文本数据的多模式Weibull变分自动编码器

