arrow
返回

Optimizing Vision Transformers: Unveiling 'Focus and Forget' for Enhanced Computational Efficiency

delete2025-01-01
delete0
delete
OA
AI
B
Banafsheh Saber Latibari *
H
Houman Homayoun
A
Avesta Sasan
DOI:10.1109/ACCESS.2025.3540399delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Vision Transformers are renowned for their accuracy in computer vision tasks but are computationally and memory expensive, making them challenging to deploy on resource-constrained edge devices. In our research paper, we introduce a revolutionary approach to designing energy-aware dynamically prunable Vision Transformers for use in edge applications. Our solution denoted as Incremental Resolution Enhancing Transformer (IRET), works by the sequential sampling of the input image. However, in our case, the embedding size of input tokens is considerably smaller than prior-art solutions. This embedding is used in the first few layers of the IRET vision transformer until a reliable attention matrix is formed. Then the attention matrix is used to sample additional information using a learnable 2D lifting scheme only for important tokens and IRET drops the tokens receiving low attention scores. Hence, as the model pays more attention to a subset of tokens for its task, its focus and resolution also increase. This incremental attention-guided sampling of input and dropping of unattended tokens allow IRET to significantly prune its computation tree on demand. By controlling the threshold for dropping unattended tokens and increasing the focus of attended ones, we can train a model that dynamically trades off complexity for accuracy. Moreover, using early exiting our model is capable of doing anytime prediction. This is especially useful for real-word energy-sensitive edge devices, where accuracy and complexity could be dynamically traded based on factors such as battery life, reliability, etc.
Keyword:
Transformers
Computational modeling
Computer vision
Head
Accuracy
Predictive models
Visualization
Discrete wavelet transforms
Computational efficiency
Spatial resolution
deep learning
pruning
vision transformer

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

University of California System 封面图
University of California System
学者数:
37.6W
论文数: 33.8W
被引数: 6.6K
引用论文

引用论文

err分享
err收藏
Measured versus label declared macronutrient and calorie content in Colombian commercially available whey proteins
err2022-07-05
err0
errOAAI
errAndrés Zapata-Muriel; Patricia Echeverry; Trisha A. Van Dusseldorp; Jennifer Kurtz; Matías Monsalves-Alvarez
err分享
err收藏
err分享
err收藏
The role of visual short-term memory in empty cell localization
err2005-11-01
err0
errOAAI
errAndrew Hollingworth; Joo-Seok Hyun; Weiwei Zhang
err分享
err收藏
err分享
err收藏
学者 查看更多内容