arrow
Return

Inter-instance similarity modeling for contrastive learning

delete2026-07-01
delete0
PRE
AI
C
Chengchao Shen *
D
Dawei Liu
H
Hao Tang
Q
Qu, Zhe
王
王健鑫 (Jianxin Wang)
DOI:10.1016/j.patcog.2026.114509delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The existing contrastive learning methods widely adopt one-hot instance discrimination as pretext task for self-supervised learning, which inevitably neglects rich inter-instance similarities among natural images, thus leading to potential representation degeneration. In this paper, we propose a novel image mix method, PatchMix, for contrastive learning in Vision Transformer (ViT), to model inter-instance similarities among images. Following the nature of ViT, we randomly mix multiple images from mini-batch in patch level to construct mixed image patch sequences for ViT. Compared to the existing sample mix methods, our PatchMix can flexibly and efficiently mix more than two images and simulate more complicated similarity relations among natural images. In this manner, our contrastive framework can significantly reduce the gap between contrastive objective and ground truth in reality. Experimental results demonstrate that our proposed method significantly outperforms the previous state-of-the-art on both ImageNet-1K and CIFAR datasets, e.g., 3.0% linear probing accuracy improvement on ImageNet-1K and 8.7% k-NN accuracy improvement on CIFAR100. Moreover, our method achieves the leading transfer performance on downstream tasks, object detection and instance segmentation on COCO dataset. The code and trained weights are available at https://github.com/ visresearch/patchmix.
Keywords:
Vision Transformer
Representation learning
Unsupervised Learning

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

C
Central South University
Scholars:
4.6K
Papers: 1.2K
Citations: 0
Cited Papers

Cited Papers

No cited papers available