arrow
Return

Coinf: QoS-aware DRL-based Inference Task Scheduling Framework with Batching Processing

delete2026-01-01
delete0
PRE
AI
张光林 cover
张光林 (Guanglin Zhang)
Y
Yuhao Zhang
X
Xiaowen Huang
W
Wenqian Zhang *
DOI:10.1145/3777373delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The emergence of deploying Deep neural network (DNN) services on edge servers has spurred research into efficiently provisioning inference services. However, previous studies have neglected to consider the implications of different types of DNN and varying quality of service (QoS) requirements on QoS violation rates. In this article, we propose a novel framework, named Coinf, for scheduling heterogeneous DNN inference tasks on edge servers. Coinf has the following four advantages to effectively handle attribute analysis, performance balancing, parallel execution, and model accuracy: (1) It enables efficient profiling of domain-specific attributes of various DNN tasks during the offline stage, achieved by constructing a regression model to predict the end-to-end latency of each task. (2) By utilizing the predicted execution time, Coinf achieves a commendable balance among inference latency, system throughput, and QoS violation rate. (3) It employs emerging deep reinforcement learning (DRL) to aggregate individual DNN tasks into batches, enabling concurrent parallel execution. (4) Coinf preserves the accuracies of the provided DNN models by not modifying them. Numerical experiments are constructed to validate the reliability and efficiency of Coinf in handling heterogeneous inference tasks.
Keywords:
Edge computing
DNN inference
deep reinforcement learning
quality of service
batching policy

Journal

ACM Transactions on Embedded Computing Systems cover
ACM Transactions on Embedded Computing Systems
IF:
2.6
Papers:
225
Citations:
2.3K

Organization

D
donghua university
Scholars:
3.3K
Papers: 972
Citations: 2