返回
Communication-Aware Supernode Shape
DOI:10.1109/TPDS.2008.114.png)
摘要
En 中文
In this paper, we revisit the supernode- shape selection problem, which has been widely discussed in the bibliography. In general, the selection of the supernode transformation greatly affects the parallel execution time of the transformed algorithm. Since the minimization of the overall parallel execution time via an appropriate supernode transformation is very difficult to accomplish, researchers have focused on scheduling-aware supernode transformations that maximize parallelism during the execution. In this paper, we argue that the communication volume of the transformed algorithm is an important criterion, and its minimization should be given high priority. For this reason, we define the metric of the per-process communication volume and propose a method to minimize this metric by selecting a communication-aware supernode shape. Our approach is equivalent to defining a proper Cartesian process grid with MPI_Cart_Create, which means that it can be incorporated in applications in a straightforward manner. Our experimental results illustrate that by selecting the tile shape with the proposed method, the total parallel execution time is significantly reduced due to the minimization of the communication volume, despite the fact that a few more parallel execution steps are required.
Keyword:
Loop tiling
supernode transformation
tile shape
MPI
process grid
scheduling
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
A unified framework for optimizing locality, parallelism, and communication in out-of-core computations用于优化核外计算中的局部性,并行性和通信的统一框架

