arrow
返回

Controlling the granularity of automatic parallel programs

delete2016-11-01
delete7
PRE
AI
A
Alcides Fonseca *
B
Bruno Cabral
DOI:10.1016/j.jocs.2016.06.005delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Programming for concurrent platforms, such as multicore cpus, is very time consuming and requires fine tuning of the final program in order to optimize the program parallel layout to the hardware architecture. Parallelization of programs is done by identification parts of code (tasks) that can be executed concurrently and execution in different threads. Current approaches for automatic parallelization cannot achieve the same performance of manually parallelized programs. Current tools are limited and either parallelize everything possible, or are limited to parallelizing the outer loops, which may miss potential parallelism that could improve the program. Some approaches have controlled granularity during execution only, but without any relevant speedups. Automatic Parallelizing Compilers have shown little overall speedup without the manual guidance of programmers in terms of granularity. This work addresses the issue of achieving performant programs from a fully automated parallelization. We propose a cost-model to decide between different parallelization alternatives. By performing static analysis, we are able to estimate the time of tasks and parallelize them only if the time is larger than the overhead of task spawning. Because the information during compilation might not be enough to make that decision, we delay some of the decisions to runtime, when all variables are available. Thus, we use an hybrid approach that performs optimizations at compile-time and at runtime. Although we apply our model in the Java language on top of the AEminium runtime, our approach is modular and can be applied to any programming language in any task-based runtime for shared-memory. We have evaluated our approach in existing benchmark programs, in cases where a wrong granularity value would result in slowing down the programs. We were able to achieve speedups greater than versions without granularity control, or with runtime-based granularity control information. We were also able to generate programs with better performance than the state-of-the-art Java automatic parallelizing compiler. Finally, in some cases we were able to outperform the human programmer. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Compiler
Parallel
Granularity
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Nature Computational Science 封面图
Nature Computational Science
IF:
18.3
论文数:
3.1K
被引数:
4.0K

机构

U
universidade de coimbra
学者数:
1.9W
论文数: 1.6W
被引数: 16
引用论文

引用论文

Exploring topological phases with quantum walks
err2010-09-24
err0
errOAAI
errTakuya Kitagawa; Mark S. Rudner; Erez Berg; Eugene Demler
err分享
err收藏
MA23.10 Low Number of Mutations and Frequent Co-Deletions of CDKN2A and IFN Type I Characterize Malignant Pleural MesotheliomaMA23.10 低突变数量和CDKN2A与I型干扰素频繁共缺失是恶性胸膜间皮瘤的特征
err2019-10-01
err0
errOAAI
errA. Nastase; A. Mandal; S.K. Lu; S. Gennatas; H. Anbunathan; M. Edwards; D. Morris-Rosendahl; A. Newman Taylor; R.C. Rintoul; E. Lim; S. Popat; A. Nicholson; M. Lathrop; A. Bowcock; M. Moffatt; W. Cookson
err分享
err收藏
没有更多内容