arrow
返回

Accelerating Deep Learning Inference via Model Parallelism and Partial Computation Offloading

delete2023-02-01
delete39
PRE
AI
H
Huan Zhou *
M
Mingze Li
N
Ning Wang *
G
Geyong Min
X
Xiangyu Wu
DOI:10.1109/TPDS.2022.3222509delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the rapid development of Internet-of-Things (IoT) and the explosive advance of deep learning, there is an urgent need to enable deep learning inference on IoT devices in Mobile Edge Computing (MEC). To address the computation limitation of IoT devices in processing complex Deep Neural Networks (DNNs), computation offloading is proposed as a promising approach. Recently, partial computation offloading is developed to dynamically adjust task assignment strategy in different channel conditions for better performance. In this paper, we take advantage of intrinsic DNN computation characteristics and propose a novel Fused-Layer-based (FL-based) DNN model parallelism method to accelerate inference. The key idea is that a DNN layer can be converted to several smaller layers in order to increase partial computation offloading flexibility, and thus further create the better computation offloading solution. However, there is a trade-off between computation offloading flexibility as well as model parallelism overhead. Then, we investigate the optimal DNN model parallelism and the corresponding scheduling and offloading strategies in partial computation offloading. In particular, we propose a Particle Swarm Optimization with Minimizing Waiting (PSOMW) method, which explores and updates the FL strategy, path scheduling strategy, and path offloading strategy to reduce time complexity and avoid invalid solutions. Finally, we validate the effectiveness of the proposed method in commonly used DNNs. The results show that the proposed method can reduce the DNN inference time by an average of 12.75 times compared to the legacy No FL (NFL) algorithm, and is very close to the optimal solution achieved by the Brute Force (BF) algorithm with the difference of less than 0.04%.
Keyword:
Mobile edge computing
fused-layer
DNN inference
partial offloading
model parallelism

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

C
china three gorges university
学者数:
1.1W
论文数: 6.1K
被引数: 114
R
Rowan University
学者数:
3.5K
论文数: 2.6K
被引数: 2.2K
U
University of Exeter
学者数:
2.0W
论文数: 2.1W
被引数: 3.6W
P
pennsylvania commonwealth system of higher education (pcshe)
学者数:
12.9W
论文数: 11.7W
被引数: 177
学者 查看更多机构
引用论文

引用论文

A Novel Mobile Edge Network Architecture with Joint Caching-Delivering and Horizontal Cooperation
err2021-01-01
err43
errOAAI
errSaputra, Yuris Mulya; Hoang, Dinh Thai; Nguyen, Diep N.; Dutkiewicz, Eryk
err分享
err收藏
Mobile Edge Computing: A Survey移动边缘计算: 一项调查
err2018-02-01
err2.0K
errOAAI
errAbbas, Nasir; Zhang, Yan; Taherkordi, Amir; Skeie, Tor
err分享
err收藏
Comprehensive learning particle swarm optimizer for global optimization of multimodal functions
err2006-06-01
err3.2K
PREAI
errLiang, J. J.; Qin, A. K.; Suganthan, Ponnuthurai Nagaratnam; Baskar, S.
err分享
err收藏
Deep Reinforcement Learning for Cooperative Content Caching in Vehicular Edge Computing and Networks
err2020-01-01
err248
PREAI
errQiao, Guanhua; Leng, Supeng; Maharjan, Sabita; Zhang, Yan; Ansari, Nirwan
err分享
err收藏
err分享
err收藏
Cooperative Task Scheduling for Computation Offloading in Vehicular Cloud
err2018-11-01
err151
PREAI
errSun, Fei; Hou, Fen; Cheng, Nan; Wang, Miao; Zhou, Haibo; Gui, Lin; Shen, Xuemin
err分享
err收藏
学者 查看更多内容