arrow
返回

Simultaneous image patch attention and pruning for patch selective transformer

delete2024-10-01
delete0
PRE
AI
S
Sunpil Kim
G
Gang-Joon Yoon
J
Jinjoo Song
S
Sang Min Yoon *
DOI:10.1016/j.imavis.2024.105239delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Vision transformer models provide superior performance compared to convolutional neural networks for various computer vision tasks but require increased computational overhead with large datasets. This paper proposes a patch selective vision transformer that effectively selects patches to reduce computational costs while simultaneously extracting global and local self-representative patch information to maintain performance. The interpatch attention in the transformer encoder emphasizes meaningful features by capturing the inter-patch relationships of features, and dynamic patch pruning is applied to the attentive patches using a learnable soft threshold that measures the maximum multi-head importance scores. The proposed patch attention and pruning method provides constraints to exploit dominant feature maps in conjunction with self-attention, thus avoiding the propagation of noisy or irrelevant information. The proposed patch-selective transformer also helps to address computer vision problems such as scale, background clutter, and partial occlusion, resulting in a lightweight and general-purpose vision transformer suitable for mobile devices.
Keyword:
Patch pruning
Patch emphasis
Attentive patch selection
Vision transformer

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

K
kookmin university
学者数:
3.0K
论文数: 3.3K
被引数: 2
引用论文

引用论文

Dual-Frequency Four-Stage Polarimetric SAR Interferometry for Forest Height Estimation
err2024-01-01
err1
PREAI
errXue, Fengli; Wang, Jili; Zheng, Mingjie; Zhang, Heng; Liu, Xiuqing; Wang, Longxiang; Jia, Xiaoxue; Deng, Yunkai
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏