arrow
Return

Multi-scale vision transformer classification model with self-supervised learning and dilated convolution

delete2022-10-01
delete2
PRE
AI
L
Liping Xing *
金红梅 cover
金红梅 (Hongmei Jin)
H
Hong-an Li
Z
Zhanli Li
DOI:10.1016/j.compeleceng.2022.108270delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Benefiting from the advantages of good parallelism and features that support long-distance dependency modeling, a variety of ViT models based on the self-attention mechanism show outstanding performance in image classification tasks. However, these works have poor classifi-cation accuracy when trained on small datasets owing to insufficient attention toward local features. Therefore, this study presents a new self-supervised multi-scale ViT classification model, SMvT. This model adopts twin-tower architectures as the self-supervised framework and the hierarchical Swin Transformer as the backbone and proposes a MDP embedding layer to fully pay attention to local details. We investigated the model's performance pretrained using the lightweight ImageNet dataset. Compared with recent self-supervised Transformers for vision, such as MoCo v3 and DINO, the classification accuracy of the SMvT has improved by up to 15.2%. SMvT combines self-supervised learning and depthwise separable dilated convolution, which is a lightweight and high generalization ViT model supporting cross-scale attention modeling.
Keywords:
Self-attention
Vision transformer
Multi-scale
Self-supervised
Depthwise separable dilated convolution
Local features

Journal

C
Computers and Electrical Engineering
IF:
4.9
Papers:
6.7K
Citations:
1.3W

Organization

X
xi'an university of science & technology
Scholars:
6.9K
Papers: 4.8K
Citations: 5