arrow
Return

RuYi: Optimizing Burst Buffer Through Automated, Fine-Grained Process-to-BB Mapping

delete2025-03-01
delete0
delete
OA
AI
Y
Yusheng Hua
X
Xuanhua Shi *
L
Ligang He
K
Kang He
张腾 (Teng Zhang)
金海 (Hai Jin)
Y
Yong Chen
DOI:10.1109/TC.2024.3510624delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Current supercomputers use an SSD-based storage layer called Burst Buffer (BB) to provide I/O-intensive applications with accelerated storage access. However, efficiently utilizing this limited and expensive storage remains a critical issue, creating an urgent need for implementing Quality of Service (QoS) in BB. To address this, we propose RuYi, a QoS-aware method to provide applications with bandwidth guarantees in the BB file system. RuYi tackles two main issues. First, it quantitatively profiles available bandwidth resources in BB to ensure reliable QoS, a crucial aspect seldom studied in the literature. Second, RuYi offers fine-grained process-level QoS via an innovative process-to-BB mapping, maximizing resource utilization-something not achievable with conventional coarse-grained compute-to-BB mapping. We evaluated RuYi on a subsystem of the leading exascale supercomputer Sunway, consisting of 4,000 compute nodes and 200 BB nodes. The experimental results demonstrate that RuYi achieves an impressive end-to-end bandwidth control accuracy of 97%, while improving BB utilization by up to 116% compared to conventional coarse-grained compute-to-BB mapping.
Keywords:
Bandwidth
Quality of service
Resource management
Interference
Supercomputers
File systems
Reliability
Processor scheduling
Degradation
Accuracy
HPC
burst buffer
file system
QoS

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

Texas Tech University System cover
Texas Tech University System
Scholars:
1.5W
Papers: 1.3W
Citations: 15
U
University of Warwick
Scholars:
2.2W
Papers: 2.2W
Citations: 85