arrow
Return

Improving Oversubscribed GPU Memory Performance in the PyTorch Framework

delete2022-11-11
delete2
PRE
AI
J
Jake Choi
H
Heon Y. Yeom
K
Kim, Yoonhee *
DOI:10.1007/s10586-022-03805-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Popular deep learning frameworks like PyTorch utilize GPUs heavily for training, and suffer from out-of-memory (OOM) problems if memory is not managed properly. CUDA Unified Memory (UM) allows the oversubscription of tensor objects in the GPU, but suffers from heavy performance penalties. In this paper, we build upon our UM implementation and create and utilize a minimal overhead CUPTI dynamic profiler to trace unified memory page fault and memory transfer statistics in PyTorch applications. We also implement CUDA memory prefetch and advise API which can be called directly from the PyTorch application based on the dynamically profiled statistics to improve oversubscription performance in various PyTorch models including Resnet and BERT.
Keywords:
CUDA
Unified memory
PyTorch
prefetch
Advise
CUPTI

Journal

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
Papers:
5.0K
Citations:
7.5K

Organization

S
seoul national university (snu)
Scholars:
7.2W
Papers: 6.6W
Citations: 86