arrow
返回

PWLT: Pyramid Window-based Lightweight Transformer for image classification

delete2024-05-01
delete2
PRE
AI
Y
Yuwei Mo
P
Pengfei Zuo
Q
Quan Zhou *
Z
Zhiyi Mo
Y
Yawen Fan
S
Suofei Zhang
B
Bin Kang
DOI:10.1016/j.compeleceng.2024.109209delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Recently, vision Transformers (ViTs) have achieved remarkable progress for image classification. As the computational cost of self -attention adopted in ViTs is quadratic with respect to the number of input tokens, some window -based ViTs have been proposed to solve this issue. However, these methods limit the computation of self -attention into spatial -constrained local windows, losing capability to encode image -based global interactions. Additionally, using fixed -size window always suffers the limitation of single -scale representation that is unsuitable for object recognition with variable scales. To address these problems, this paper describes a Pyramid Window -based Lightweight Transformer, namely PWLT, for image classification. Specifically, to address the need for multi -scale information, we employ windows of different sizes to encode objects with varying scales. To restore the relationships between different windows and explore global context, we introduce a dual self -attention scheme that utilizes local -to -global attention to reestablish these relationships. The extensive experiments on ImageNet-1K and CIFAR100 datasets demonstrate the effectiveness of our PWLT for image classification.
Keyword:
Image classification
Lightweight vision transformer
Pyramid window
Self-attention

期刊

C
Computers and Electrical Engineering
IF:
4.9
论文数:
6.7K
被引数:
1.3W

机构

W
Wuzhou University
学者数:
324
论文数: 310
被引数: 454