Return
Prediction-based GPU Sharing for Distributed Training
DOI:10.1016/j.future.2026.108413.png)
Abstract
En 中文
• Formulate the inconsistent JCT problem using gSLA for the first time. • Design a new JCT increase prediction model and job scheduler for GPU sharing. • Achieve up to 47.3× better gSLA satisfaction and 50× lower gSLA excess ratio. • Improve JCT and GPU efficiency by ∼ 60% and ∼ 44% over existing methods. • Demonstrate TensorShare’s effectiveness in improving gSLA and JCT for unseen jobs.
Keywords:
Cloud computing
GPU sharing
Service level agreement
Performance prediction
GPU scheduling
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
F
IF:
0
Papers:
642
Citations:
0
Organization
Cited Papers
Managing Performance Overhead of Virtual Machines in Cloud Computing: A Survey, State of the Art, and Future Directions
PROCEEDINGS OF THE IEEE
IF25.9

