返回
Robust Contrastive Learning With Dynamic Mixed Margin
DOI:10.1109/ACCESS.2023.3286931.png)
摘要
En 中文
One of the promising ways for the representation learning is contrastive learning. It enforces that positive pairs become close while negative pairs become far. Contrastive learning utilizes the relative proximity or distance between positive and negative pairs. However, contrastive learning might fail to handle the easily distinguished positive-negative pairs because the gradient of easily divided positive-negative pairs comes to vanish. To overcome the problem, we propose a dynamic mixed margin (DMM) loss that generates the augmented hard positive-negative pairs that are not easily clarified. DMM generates hard positive-negative pairs by interpolating the dataset with Mixup. Besides, DMM adopts the dynamic margin incorporating the interpolation ratio, and dynamic adaptation improves representation learning. DMM encourages making close for positive pairs far away, whereas making a little far for strongly nearby positive pairs alleviates overfitting. Our proposed DMM is a plug-and-play module compatible with diverse contrastive learning loss and metric learning. We validate that the DMM is superior to other baselines on various tasks, video-text retrieval, and recommender system task in unimodal and multimodal settings. Besides, representation learned from DMM shows better robustness even if the modality missing occurs that frequently appears on the real-world dataset. Implementation of DMM at downstream tasks is available here: https://github.com/teang1995/DMM
Keyword:
Task analysis
Representation learning
Measurement
Business process re-engineering
Visualization
Transformers
Face recognition
Multimodal learning
contrastive learning
retrieval
video representation
recommender system
robustness
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Conjugated heat transfer and temperature distributions in a gas turbine combustion liner under base-load operation基本负荷运行下燃气轮机燃烧衬里中的共轭传热和温度分布
Multimodal Sentiment Intensity Analysis in Videos: Facial Gestures and Verbal Messages视频中的多模态情感强度分析: 面部手势和言语信息

