Return
EasyViT: An Adaptive Collaborative Edge Computing Framework for Vision Transformer
DOI:10.1109/JIOT.2025.3578605.png)
Abstract
En 中文
Deploying vision transformers (ViTs) in edge computing environments presents significant challenges due to their high computing demands and the resource constraints of edge devices. While collaborative edge computing and dynamic token dropping offer potential solutions, existing approaches suffer from rigid strategies that fail to adapt to diverse conditions of network and computing resources at the edge. This article introduces EasyViT, an adaptive framework that optimizes ViT deployment through the joint coordination of collaborative edge computing and dynamic token dropping. Key innovations include: 1) A token dropping model that integrates dynamic token dropping and collaborative edge environments, formulating an integer linear programming (ILP) optimization problem. 2) An approximate stochastic gradient descent (ASGD) method with atomic gradient calculation, which transforms the NP-hard ILP problem into a continuous space for rapid near-optimal solution generation. Extensive evaluations on a real-world edge testbed with multiple Raspberry Pi nodes demonstrate that EasyViT achieves 1.06– $5.06\times $ speedup over baseline methods under 20 configurations of edge environments, while maintaining model accuracy within 2.8% degradation. The proposed framework exhibits the adaptability across diverse ViT architectures, network bandwidths, and computing resources.
Keywords:
Collaborative computing
edge computing
self-attention
Vision Transformer
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

