arrow
Return

Resource-Efficient Orchestration for Heterogeneous Serverless Computing with Harmonized Effectiveness and Practicability

delete2026-05-01
delete1
PRE
AI
L
Liu, Qingyuan
Y
Yang, Yanning
D
Du, Dong *
Z
Zhang, Ping
F
Feng, Jia
C
Chen, Haibo
DOI:10.1145/3788863delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Current serverless platforms struggle to optimize resource utilization for both CPU and GPU functions due to their dynamic and fine-grained nature. Conventional techniques like overcommitment and autoscaling fall short, often sacrificing utilization for practicability or incurring performance tradeoffs. Overcommitment requires predicting performance to prevent QoS violation, introducing tradeoff between prediction accuracy and overheads. Autoscaling requires scaling instances in response to load fluctuations quickly to reduce resource wastage, but more frequent scaling also leads to more cold start overheads. The rich concurrency of GPU resources further complicates GPU instance orchestration, such as setting right batch sizes. This article introduces JIAGU to harmonize efficiency with practicability through the following novel techniques. First, pre-decision scheduling achieves accurate prediction while eliminating overheads by decoupling prediction and scheduling. Second, dual-staged scaling achieves frequent adjustment of instances with minimum overhead. Third, JIAGU conducts an in-depth analysis about the complexity of the relationship between GPU function configuration and execution. It then proposes batch-aware scaling that achieves optimal configurations for both batch size setting and autoscaling, addressing all the challenges according to the analysis. We have implemented a prototype and evaluated it using real-world applications and traces from the public cloud platform. Our evaluation shows an improvement in deployment density over commercial clouds (with Kubernetes) while maintaining QoS for both CPU and GPU functions (54.8% and 18% respectively), and 81.0%-93.7% lower scheduling costs and a 57.4%-69.3% reduction in cold start latency compared to existing QoS-aware schedulers.
Keywords:
Serverless computing
instance scheduling
autoscaling
batch inference

Journal

A
ACM Transactions on Computer Systems
IF:
1.8
Papers:
26
Citations:
1.1K

Organization

S
shanghai jiao tong university
Scholars:
15.5W
Papers: 11.6W
Citations: 159