Return
Data-efficient multi-scale fusion vision transformer
DOI:10.1016/j.patcog.2024.111305.png)
Abstract
En 中文
Vision transformers (ViTs) excel in image classification with large datasets but struggle with smaller ones. Vanilla ViTs are single-scale, tokenizing images into patches with a single patch size. In this paper, we introduce multi-scale tokens, where multiple scales are achieved by splitting images into patches of varying sizes. Our model concatenates token sequences of multiple scales for attention, and a regional cross-scale interaction module fuses these tokens, improving data efficiency by learning local structures across scales. Additionally, we implement a data augmentation schedule to refine training. Extensive experiments on image classification demonstrate our approach surpasses DeiT by 6.6% on CIFAR100 and 1.6% on ImageNet1K. Code is available at https://github.com/visresearch/dems.
Keywords:
Deep learning
Image classification
Vision transformer
Data efficiency
Multi-scale fusion
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W
Organization
No organization information available

