arrow
Return

Simultaneous image patch attention and pruning for patch selective transformer

delete2024-10-01
delete0
PRE
AI
S
Sunpil Kim
G
Gang-Joon Yoon
J
Jinjoo Song
S
Sang Min Yoon *
DOI:10.1016/j.imavis.2024.105239delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Vision transformer models provide superior performance compared to convolutional neural networks for various computer vision tasks but require increased computational overhead with large datasets. This paper proposes a patch selective vision transformer that effectively selects patches to reduce computational costs while simultaneously extracting global and local self-representative patch information to maintain performance. The interpatch attention in the transformer encoder emphasizes meaningful features by capturing the inter-patch relationships of features, and dynamic patch pruning is applied to the attentive patches using a learnable soft threshold that measures the maximum multi-head importance scores. The proposed patch attention and pruning method provides constraints to exploit dominant feature maps in conjunction with self-attention, thus avoiding the propagation of noisy or irrelevant information. The proposed patch-selective transformer also helps to address computer vision problems such as scale, background clutter, and partial occlusion, resulting in a lightweight and general-purpose vision transformer suitable for mobile devices.
Keywords:
Patch pruning
Patch emphasis
Attentive patch selection
Vision transformer

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

K
kookmin university
Scholars:
3.0K
Papers: 3.3K
Citations: 2