arrow
Return

Dynamic Batch Processing with FlexiDecode Scheduler for Efficient LLM Inference in IIoT

delete2025-12-01
delete0
PRE
AI
X
Xiaocong Jia
B
Bruce Gu
J
Jinjun Chen
L
Longxiang Gao *
W
Weiguang Pang
G
G.D. Lv
Y
Youyang Qu
L
Lei Cui
DOI:10.26599/BDMA.2025.9020025delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large Language Models (LLMs) are expanding their applications across various fields, including Industrial Internet of Things (IIoT), where they analyze sensor data, automate diagnostics, and enhance predictive maintenance. LLM inference is provided by service providers to users, with each inference request undergoing two phases: prefill and decode. Due to the autoregressive nature of generation, only one token can be produced per iteration, necessitating multiple iterations to complete a request. Typically, batch processing groups multiple requests into a single batch for inference, improving throughput and hardware utilization. However, in service systems, a fixed batch size presents challenges under fluctuating request volumes, particularly in IIoT environments, where data flow can vary significantly. Specifically, during the high-load periods, a fixed batch size may lead to underutilization of resources, while during the low-load periods, it may result in resource wastage. In this paper, we introduce FlexiDecode Scheduler (FDS) to address these challenges by dynamically adjusting the decoding batch size based on system load conditions, improving resource utilization, and reducing wait time during high-load periods. FDS prioritizes prefilling new requests to maximize decoding efficiency and employs a request output length predictor to optimize request scheduling, minimizing End-to-End (E2E) latency. Compared to virtual Large Language Model (vLLM) and Sarathi, our approach achieves a 23% and 16% reduction in E2E latency, improves actual request execution time by 34% and 15%, respectively, and increases computational utilization by 10%.
Keywords:
virtual Large Language Model (vLLM) inference
batch scheduling
batch scheduling
dynamic decoding batches
dynamic decoding batches
calculating utilization
calculating utilization
calculating utilization

Journal

Big Data Mining and Analytics cover
Big Data Mining and Analytics
IF:
6.2
Papers:
274
Citations:
1.0K

Organization

Q
Qilu University of Technology
Scholars:
1.1W
Papers: 8.9K
Citations: 16