arrow
返回

Optimizing process creation and execution on multi-core architectures

delete2013-04-02
delete1
PRE
AI
K
Kulkarni, Abhishek *
L
Latchesar Ionkov
M
Michael Lang
A
Andrew Lumsdaine
DOI:10.1177/1094342013481483delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The execution of a single process multiple data (SPMD) application involves running multiple instances of a process with possibly varying arguments. With the widespread adoption of massively multicore processors, there has been a focus towards harnessing the abundant compute resources effectively in a power-efficient manner. Although much work has been done towards optimizing distributed process launch using hierarchical techniques, there has been a void in studying the performance of spawning processes within a single node. Reducing the latency to spawn a new process locally results in faster global job launch. Further, emerging dynamic and resilient execution models are designed on the premise of maintaining process pools for fault isolation and launching several processes in a relatively shorter period of time. Optimizing the latency and throughput for spawning processes would help improve the overall performance of runtime systems, allow adaptive process-replication reliability and motivate the design and implementation of process management interfaces in future manycore operating systems. In this paper, we study the several limiting factors for efficient spawning of processes on massively multicore architectures. We have developed a library to optimize launching multiple instances of the same executable. Our microbenchmarks show a 20-80% decrease in the process spawn time for multiple executables. We further discuss the effects of memory locality and propose NUMA-aware extensions to optimize launching processes with large memory-mapped segments including dynamic shared libraries. Finally, we describe vector operating system interfaces for spawning a batch of processes from a given executable on specific cores. Our results show a speedup of a factor of 40-50 over the traditional method of launching new processes using fork and exec system calls.
Keyword:
process spawn optimization
scalable intra-node process launch
batched system calls
vector operating system interfaces
manycore runtime systems
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

International Journal of High Performance Computing Applications 封面图
International Journal of High Performance Computing Applications
IF:
2.5
论文数:
1.1K
被引数:
1.3K

机构

I
indiana university system
学者数:
4.0W
论文数: 3.5W
被引数: 38
I
Indiana University Bloomington
学者数:
1.9W
论文数: 1.5W
被引数: 2.8W
引用论文

引用论文

Simulation and Optimization of Aircraft Assembly Process Using Supercomputer Technologies
err2018-12-31
err0
PREAI
errTatiana Pogarskaia; Maria Churilova; Margarita Petukhova; Evgeniy Petukhov
err分享
err收藏
In situ phase change characterization of PVDF thin films using Raman spectroscopy
err2014-04-10
err0
errOAAI
errMiranda T. Riosbaas; Kenneth J. Loh; Greg O'Bryan; Bryan R. Loyola
err分享
err收藏
err分享
err收藏
没有更多内容