Return
Multi-scale vision transformer classification model with self-supervised learning and dilated convolution
DOI:10.1016/j.compeleceng.2022.108270.png)
Abstract
En 中文
Benefiting from the advantages of good parallelism and features that support long-distance dependency modeling, a variety of ViT models based on the self-attention mechanism show outstanding performance in image classification tasks. However, these works have poor classifi-cation accuracy when trained on small datasets owing to insufficient attention toward local features. Therefore, this study presents a new self-supervised multi-scale ViT classification model, SMvT. This model adopts twin-tower architectures as the self-supervised framework and the hierarchical Swin Transformer as the backbone and proposes a MDP embedding layer to fully pay attention to local details. We investigated the model's performance pretrained using the lightweight ImageNet dataset. Compared with recent self-supervised Transformers for vision, such as MoCo v3 and DINO, the classification accuracy of the SMvT has improved by up to 15.2%. SMvT combines self-supervised learning and depthwise separable dilated convolution, which is a lightweight and high generalization ViT model supporting cross-scale attention modeling.
Keywords:
Self-attention
Vision transformer
Multi-scale
Self-supervised
Depthwise separable dilated convolution
Local features
Journal
C
IF:
4.9
Papers:
6.7K
Citations:
1.3W

