arrow
Return

SemDM: Task-oriented masking strategy for self-supervised visual learning

delete2023-09-01
delete0
PRE
AI
X
Xin Ma
H
Haonan Cheng
叶龙 cover
叶龙 (Long Ye) *
DOI:10.1016/j.displa.2023.102439delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper, we propose a novel learning scheme for better combining masked image modeling (MIM) and instance discrimination (ID). Motivated by compensating the requirement gap of masking strength between MIM and ID, we propose Semantic Disjoint Masking (SemDM), which decomposes the masking into two manners: preserving the majority of key patterns in images for ID, while dropping out most of them for MIM. Specifically, we utilize attention-guided masking in ID to help keeping the identity of object in image for encoder. While in MIM, we conversely only leave some hints about the object. Then these generated masked views only perform their specified learning task, facilitating more suitable visual priors to be learned in each learning task. Moreover, we introduce product quantization (PQ) to optimize the concept distributions in latent space, which guarantees that a compact set of meaningful visual concepts can be learned. Extensive experiments demonstrate that our method bootstraps meaningful visual concepts to guide visual understanding, and obtains state-of-the-art results on ImageNet-100.
Keywords:
Masked image modeling
Masking instance discrimination
Multi-task Learning

Journal

Displays cover
Displays
IF:
3.4
Papers:
2.1K
Citations:
3.2K

Organization

C
Communication University of China
Scholars:
1.1K
Papers: 820
Citations: 326