1
Return

Teaching Masked Autoencoder With Strong Augmentations

delete2024-01-01
delete0
PRE
AI
R
Rui Zhu
Y
Yalong Bai
T
Ting Yao *
J
Jingen Liu
Z
Zhenglong Sun *
T
Tao Mei
C
Chang Wen Chen
DOI:10.1109/TNNLS.2024.3419898delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Masked autoencoder (MAE) has been regarded as a capable self-supervised learner for various downstream tasks. Nevertheless, the model still lacks high-level discriminability, which results in poor linear probing performance. In view of the fact that strong augmentation plays an essential role in contrastive learning, can we capitalize on strong augmentation in MAE? The difficulty originates from the pixel uncertainty caused by strong augmentation that may affect the reconstruction, and thus, directly introducing strong augmentation into MAE often hurts the performance. In this article, we delve into the potential of strong augmented views to enhance MAE while maintaining MAE's advantages. To this end, we propose a simple yet effective masked Siamese autoencoder (MSA) model, which consists of a student branch and a teacher branch. The student branch derives MAE's advanced architecture, and the teacher branch treats the unmasked strong view as an exemplary teacher to impose high-level discrimination onto the student branch. We demonstrate that our MSA can improve the model's spatial perception capability and, therefore, globally favors interimage discrimination. Empirical evidence shows that the model pretrained by MSA provides superior performances across different downstream tasks. Notably, linear probing performance on frozen features extracted from MSA leads to 6.1% gains over MAE on ImageNet-1k. Fine-tuning (FT) the network on VQAv2 task finally achieves 67.4% accuracy, outperforming 1.6% of the supervised method DeiT and 1.2% of MAE.
Keywords:
Task analysis
Representation learning
Image reconstruction
Decoding
Contrastive learning
Training
Semantics
data augmentation
deep learning
masked image modeling
self-supervised representation learning

Journal

IEEE Transactions on Neural Networks and Learning Systems cover
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
Papers:
7.5K
Citations:
7.2W

Organization

H
hong kong polytechnic university
Scholars:
3.0W
Papers: 4.0W
Citations: 921
T
The Chinese University of Hong Kong, Shenzhen
Scholars:
4.2K
Papers: 3.9K
Citations: 7
A
amazon.com
Scholars:
690
Papers: 500
Citations: 8
Cited Papers

Cited Papers

Citing Papers

Citing Papers