返回
Multi-Exit DNN Inference Acceleration Based on Multi-Dimensional Optimization for Edge Intelligence
DOI:10.1109/TMC.2022.3172402.png)
摘要
En 中文
Edge intelligence, as a prospective paradigm for accelerating DNN inference, is mostly implemented by model partitioning which inevitably incurs the large transmission overhead of DNN's intermediate data. A popular solution introduces multi-exit DNNs to reduce latency by enabling early exits. However, existing work ignores the correlation between exit settings and synergistic inference, causing incoordination of device-to-edge. To address this issue, this paper first investigates the bottlenecks of executing multi-exit DNNs in edge computing and builds a novel model for inference acceleration with exit selection, model partition, and resource allocation. To tackle the intractable coupling subproblems, we propose a Multi-exit DNN inference Acceleration framework based on Multi-dimensional Optimization (MAMO). In MAMO, the exit selection subproblem is first extracted from the original problem. Then, bidirectional dynamic programming is employed to determine the optimal exit setting for an arbitrary multi-exit DNN. Finally, based on the optimal exit setting, a DRL-based policy is developed to learn joint decisions of model partition and resource allocation. We deploy MAMO on a real-world testbed and evaluate its performance in various scenarios. Extensive experiments show that it can adapt to heterogeneous tasks and dynamic networks, and accelerate DNN inference by up to 13:7x compared with the state-of-the-art.
Keyword:
Edge intelligence
exit selection
inference acceleration
model partition
multi-exit DNN
resource allocation
期刊
IF:
9.2
论文数:
5.8K
被引数:
1.8W
机构
引用论文
Joint Multiuser DNN Partitioning and Computational Resource Allocation for Collaborative Edge Intelligence面向协作边缘智能的联合多用户DNN划分和计算资源分配
Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing边缘智能: 用边缘计算铺平人工智能的最后一英里
PROCEEDINGS OF THE IEEE
IF25.9
In-Edge AI: Intelligentizing Mobile Edge Computing, Caching and Communication by Federated LearningIn-Edge AI: 通过联合学习实现移动边缘计算、缓存和通信的智能化
IEEE NETWORK
IF6.3

