arrow
Return

Latency-Based Inter-Operator Scheduling for CNN Inference Acceleration on GPU

delete2024-01-01
delete0
PRE
AI
Y
Yukai Ping
H
He Jiang *
X
Xingxiang Liu
Z
Zhenyang Zhao
Z
Zhide Zhou
X
Xin Chen
DOI:10.1109/TSC.2023.3345952delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Convolutional Neural Networks (CNNs) are widely deployed on the Graphics Processing Unit (GPU) to support Deep Learning (DL) based services. Popular DL frameworks usually ignore the inter-operator parallelism when executing the inference of CNNs, which results in high inference latency. Although some inter-operator scheduling methods have been proposed, there remains a critical trade-off issue between inference latency (effectiveness) and scheduling time (efficiency). In this article, we propose LIOS, a novel latency-based heuristic inter-operator scheduling method to balance inference latency and scheduling time. In LIOS, a CNN latency model is built based on the given CNN and GPU. Then every operator is assigned a priority value to represent its importance. During each iteration of the scheduling process, LIOS identifies the current data-independent operators, selects the operator with the highest priority value, and assigns it to the GPU stream with the smallest finish time. Extensive experimental results have demonstrated the effectiveness and efficiency of LIOS. For the effectiveness, LIOS can speed up the inference of normal-size and large-size CNNs by 1.13 similar to 1.59x similar to 1.59x compared to sequential scheduling. This result is comparable to IOS, the latest state-of-the-art scheduling method. For the efficiency, LIOS can speed up the scheduling process by 7 similar to 9210x similar to 9210x compared to IOS. Like what you're reading?
Keywords:
Convolutional neural networks
Tensors
Convoluitonal neural network
deep learning
inference acceleration
inter-operator scheduling

Journal

IEEE Transactions on Services Computing cover
IEEE Transactions on Services Computing
IF:
5.8
Papers:
2.1K
Citations:
6.5K

Organization

D
Dalian University of Technology
Scholars:
5.9W
Papers: 4.4W
Citations: 5.5W