arrow
Return

Database Perspective on LLM Inference Systems

delete2025-08-01
delete0
PRE
AI
J
James Pan *
G
Guoliang Li
DOI:10.14778/3750601.3750703delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) are powering a new wave of language-based applications, including database applications, leading to new techniques and systems for dealing with the enormous compute and memory needs of LLMs, coupled with advances in computing hardware. In this tutorial, we review how these techniques lower inference costs by managing uncertain request lifecycles, exploiting specialized hardware, and scaling over distributed inference devices and machines. We present these techniques from the database perspective of request processing, model execution and optimization, and memory management. Following these discussion, we review how inference systems combine these techniques in diverse architectures to achieve application or performance objectives.

Journal

P
Proceedings of the VLDB Endowment
IF:
3.3
Papers:
563
Citations:
1.2W

Organization

T
Tsinghua University
Scholars:
8.6K
Papers: 4.1K
Citations: 17.7W
Cited Papers

Cited Papers

AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
err2025-01-21
err0
PREAI
errJi Lin; Jiaming Tang; Haotian Tang; Shang Yang; Guangxuan Xiao; Song Han
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
LLM for Data Management
err2024-11-08
err1
PREAI
errZhou, Xuanhe; Zhao, Xinyang
errShare
errSave
errShare
errSave