arrow
Return

Horizontal pod splitting scheduling for cost-effectiveness in cloud computing

delete2026-06-25
delete0
delete
OA
AI
S
Shengyi Wang *
X
Xianhui Liu
J
Jianwei Fu
S
Shunzhang Liu
X
Xiaoyong Tian
X
Xiaowei Du
DOI:10.1186/s13677-026-00945-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Optimizing service latency while controlling resource costs is a critical challenge in cloud-native microservice deployments. This paper proposes Horizontal Pod Splitting (HPS), a novel scheduling method that jointly optimizes the number of pod replicas and resource quotas under fixed total resources to minimize service latency. We construct a double-exponential service latency model that characterizes the relationships among total resources, unit resource concurrency, pod replicas, and service latency. Based on this model, we theoretically derive the optimal number of pod replicas and schedulability condition. Experimental results on a Kubernetes cluster demonstrate that HPS reduces service latency by 6–12% compared to HPA and by 7–22% compared with Libra under equivalent resource conditions, while achieving higher cloud resource utilization and reducing deployment costs for service providers.
Keywords:
Horizontal pod splitting
Resource scheduling
Service latency Model
Optimal pod replicas
Cloud computing

Journal

J
Journal of Cloud Computing-Advances Systems and Applications
IF:
4.3
Papers:
731
Citations:
2.2K

Organization

C
college of electronic and information engineering
Scholars:
163
Papers: 62
Citations: 0