arrow
Return

HASE: Hardware-Aware Scheduling for Inference Tasks in Heterogeneous GPU Clusters

delete2026-04-20
delete0
delete
OA
AI
Y
Yanqi Chen
C
Congfeng Jiang
C
Chunpeng Wu
Y
Yue Wang
Q
Qinghe Ye
L
Longchuan Yan
J
Jianing Niu
J
J J Liu
L
Lingjia Lao
DOI:10.1109/tcc.2026.3685862delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
When processing large-scale co-located inference workloads in heterogeneous GPU clusters, existing cluster scheduling mechanisms often increase the job makespan due to neglecting the mutual performance interference and resource contention for co-located tasks. This deficiency is dramatically amplified in typical neural networks based deep learning workloads because these workloads are highly sensitive to hardware configurations like GPU memory capacity, bandwidth and GPU core frequency. Therefore, a hardware-performance-aware scheduler capable of adaptively and dynamically dispatching tasks according to GPU hardware characteristics is crucial for reducing the overall completion time of job queues. To address this issue, we propose Hardware-Aware Scheduling for Inference Tasks in Heterogeneous GPU Clusters (<i>HASE</i>), a novel scheduling framework that dynamically adapts to GPU hardware characteristics for real-time optimal task placement. <i>HASE</i> consists of two core components: a kernel-level latency prediction model and a hybrid scheduling strategy. Unlike conventional approaches that rely on coarse-grained model-level features, our predictor decomposes inference models into fine-grained computational kernels using ONNX Runtime graph optimization, and predicts individual kernel execution times under varying GPU load conditions. By integrating static hardware specifications, dynamic microbenchmarks, and real-time DCGM profiling metrics, the predictor captures both operator-level heterogeneity and background load interference. The hybrid scheduling strategy combines a two-stage greedy search for task placement with a resource reservation and backfilling mechanism to balance immediate optimization and long-term fairness. Experimental results on CNN-based inference workloads (YOLO, ResNet, VGG, DenseNet, and MobileNet series) demonstrate that <i>HASE</i> achieves a kernel-level prediction accuracy of 8.2% MAPE and model-level accuracy of <inline-formula><tex-math notation="LaTeX">$R^{2}$</tex-math></inline-formula> = 0.91. Moreover, <i>HASE</i> reduces 51% total job makespan compared to traditional round-robin scheduling, and maintains per-task scheduling decision time under one second in clusters with up to hundreds of GPUs.
Keywords:
Task scheduling
heterogeneous GPU cluster
deep learning
performance prediction
inference task

Journal

I
IEEE Transactions on Cloud Computing
IF:
5
Papers:
1.8K
Citations:
4.3K

Organization

H
Hangzhou Dianzi University
Scholars:
1.3W
Papers: 9.5K
Citations: 7.5K
S
state grid
Scholars:
158
Papers: 86
Citations: 0
C
China Electric Power Research Institute
Scholars:
785
Papers: 414
Citations: 1.4K
researcher View more organizations