返回
Multi-scale vision transformer classification model with self-supervised learning and dilated convolution
DOI:10.1016/j.compeleceng.2022.108270.png)
摘要
En 中文
Benefiting from the advantages of good parallelism and features that support long-distance dependency modeling, a variety of ViT models based on the self-attention mechanism show outstanding performance in image classification tasks. However, these works have poor classifi-cation accuracy when trained on small datasets owing to insufficient attention toward local features. Therefore, this study presents a new self-supervised multi-scale ViT classification model, SMvT. This model adopts twin-tower architectures as the self-supervised framework and the hierarchical Swin Transformer as the backbone and proposes a MDP embedding layer to fully pay attention to local details. We investigated the model's performance pretrained using the lightweight ImageNet dataset. Compared with recent self-supervised Transformers for vision, such as MoCo v3 and DINO, the classification accuracy of the SMvT has improved by up to 15.2%. SMvT combines self-supervised learning and depthwise separable dilated convolution, which is a lightweight and high generalization ViT model supporting cross-scale attention modeling.
Keyword:
Self-attention
Vision transformer
Multi-scale
Self-supervised
Depthwise separable dilated convolution
Local features
期刊
C
IF:
4.9
论文数:
6.7K
被引数:
1.3W
机构
引用论文
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?

