arrow
Return

Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving

delete2026-01-01
delete0
PRE
AI
J
Jianxiong Liao
Q
Quanxing Dong
Y
Yunkai Liang
Z
Zhi Zhou
X
Xu Chen *
DOI:10.1145/3774934.3786413delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Engaging applications with diverse SLO requirements has become indispensable for production-scale LLM serving systems. However, existing systems rely on iteration-level scheduling, which enforces inflexible, unified execution across multi-SLO workloads, significantly constraining the serving efficiency. In this paper, we introduce layer-level scheduling, a novel mechanism that advances beyond conventional iterationlevel granularity. This mechanism decomposes per-iteration computation into fine-grained layer operations, enabling the tailored execution of requests with differing requirements. However, this increased granularity introduces new challenges in both intra-instance request execution and crossinstance coordination, posing significant barriers to practical deployment. To address these challenges, we introduce Laser, a system designed for efficient multi-SLO LLM serving. The key aspect lies in the seamless integration of inter-instance request dispatching with layer-level scheduling within instances, delivering high serving throughput with SLO guarantees. Evaluations with real-world applications reveal that Laser effectively improves throughput by over 1.67x while maintaining the same SLO attainment rate compared to stateof-the-art systems.
Keywords:
Layer-level Scheduling
Multi-SLO Serving
Large Language Model

Journal

P
PROCEEDINGS OF THE 31ST ACM SIGPLAN ANNUAL SYMPOSIUM ON PRINCIPLES AND PRACTICE OF PARALLEL PROGRAMMING, PPOPP 2026
IF:
0
Papers:
43
Citations:
0

Organization

S
Sun Yat sen University
Scholars:
5.9K
Papers: 1.6K
Citations: 1.8W