返回
Block size estimation for data partitioning in HPC applications using machine learning techniques
DOI:10.1186/s40537-023-00862-w.png)
摘要
En 中文
The extensive use of HPC infrastructures and frameworks for running data-intensive applications has led to a growing interest in data partitioning techniques and strategies. In fact, application performance can be heavily affected by how data are partitioned, which in turn depends on the selected size for data blocks, i.e. the block size. Therefore, finding an effective partitioning, i.e. a suitable block size, is a key strategy to speed-up parallel data-intensive applications and increase scalability. This paper describes a methodology, namely BLEST-ML (BLock size ESTimation through Machine Learning), for block size estimation that relies on supervised machine learning techniques. The proposed methodology was evaluated by designing an implementation tailored to dislib, a distributed computing library highly focused on machine learning algorithms built on top of the PyCOMPSs framework. We assessed the effectiveness of the provided implementation through an extensive experimental evaluation considering different algorithms from dislib, datasets, and infrastructures, including the MareNostrum 4 supercomputer. The results we obtained show the ability of BLEST-ML to efficiently determine a suitable way to split a given dataset, thus providing a proof of its applicability to enable the efficient execution of data-parallel applications in high performance environments.
Keyword:
Data partitioning
High performance computing
Data-parallel applications
Machine learning
Big data
期刊
IF:
6.4
论文数:
1.5K
被引数:
1.1W
机构
引用论文
Genomic DNA sequencing by SPEL-6 primer walking using hexamer ligation1Published in conjunction with A Wisconsin Gathering Honoring Waclaw Szybalski on the occasion of his 75th year and 20years of Editorship-in-Chief of Gene, 10–11 August 1997, University of Wisconsin, Madison, WI, USA.1
Gene
IF0
A Data-Aware Scheduling Strategy for Executing Large-Scale Distributed Workflows用于执行大规模分布式工作流的数据感知调度策略
IEEE ACCESS
IF3.6
Gradient-based learning applied to document recognition基于梯度的学习在文档识别中的应用
PROCEEDINGS OF THE IEEE
IF25.9

