arrow
Return

Enable importance-aware model cacheability for inference serving

delete2025-08-09
delete0
PRE
AI
H
Hao Mo
D
Didier El Baz
L
Ligu Zhu
S
Suping Wang
S
Songfu Tan
H
Hongning Zhao
L
Lei Shi *
DOI:10.1016/j.engappai.2025.111908delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Inference serving systems are leveraged to deploy deep learning (DL) models as services. Accelerators such as Graphics Processing Units (GPUs) have been extensively used in these systems to reduce model execution time. As accelerators become more powerful and expensive, GPU sharing among DL models across different inference requests is a common practice. However, GPU memory capacity becomes a bottleneck when the number of collocated models increases, making this approach unsustainable. At the same time, collocated models may vary in popularity levels — some are accessed frequently and others are not, leading to low resource efficiency and system performance. While some existing inference serving systems offer the capability to dynamically load and cache model in memory, they are typically locality-aware and exhibit poor performance for DL inference serving.
Keywords:
GPU sharing
deep learning models
inference serving
GPU memory capacity
dynamic loading

Journal

Engineering Applications of Artificial Intelligence cover
Engineering Applications of Artificial Intelligence
IF:
8
Papers:
5.3K
Citations:
3.5W

Organization

C
Communication University of China
Scholars:
1.1K
Papers: 820
Citations: 326
Université de Toulouse cover
Université de Toulouse
Scholars:
1.4K
Papers: 593
Citations: 2.3W