Return
Enable importance-aware model cacheability for inference serving
DOI:10.1016/j.engappai.2025.111908.png)
Abstract
En 中文
Inference serving systems are leveraged to deploy deep learning (DL) models as services. Accelerators such as Graphics Processing Units (GPUs) have been extensively used in these systems to reduce model execution time. As accelerators become more powerful and expensive, GPU sharing among DL models across different inference requests is a common practice. However, GPU memory capacity becomes a bottleneck when the number of collocated models increases, making this approach unsustainable. At the same time, collocated models may vary in popularity levels — some are accessed frequently and others are not, leading to low resource efficiency and system performance. While some existing inference serving systems offer the capability to dynamically load and cache model in memory, they are typically locality-aware and exhibit poor performance for DL inference serving.
Keywords:
GPU sharing
deep learning models
inference serving
GPU memory capacity
dynamic loading
Journal
IF:
8
Papers:
5.3K
Citations:
3.5W

