arrow
Return

EasyViT: An Adaptive Collaborative Edge Computing Framework for Vision Transformer

delete2025-08-12
delete0
PRE
AI
D
Dong Wen
G
Guanping Liang
李天匀 (Tianyun Li)
L
Lin Chen
J
Junnan Li
T
Tao Li
DOI:10.1109/JIOT.2025.3578605delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deploying vision transformers (ViTs) in edge computing environments presents significant challenges due to their high computing demands and the resource constraints of edge devices. While collaborative edge computing and dynamic token dropping offer potential solutions, existing approaches suffer from rigid strategies that fail to adapt to diverse conditions of network and computing resources at the edge. This article introduces EasyViT, an adaptive framework that optimizes ViT deployment through the joint coordination of collaborative edge computing and dynamic token dropping. Key innovations include: 1) A token dropping model that integrates dynamic token dropping and collaborative edge environments, formulating an integer linear programming (ILP) optimization problem. 2) An approximate stochastic gradient descent (ASGD) method with atomic gradient calculation, which transforms the NP-hard ILP problem into a continuous space for rapid near-optimal solution generation. Extensive evaluations on a real-world edge testbed with multiple Raspberry Pi nodes demonstrate that EasyViT achieves 1.06– $5.06\times $ speedup over baseline methods under 20 configurations of edge environments, while maintaining model accuracy within 2.8% degradation. The proposed framework exhibits the adaptability across diverse ViT architectures, network bandwidths, and computing resources.
Keywords:
Collaborative computing
edge computing
self-attention
Vision Transformer

Journal

IEEE Internet of Things Journal cover
IEEE Internet of Things Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

Organization

N
National University of Defense Technology
Scholars:
3.3K
Papers: 1.0K
Citations: 8.2K