返回
Error detection in large-scale parallel programs with long runtimes
DOI:10.1016/S0167-739X(02)00178-4.png)
摘要
En 中文
Error detection is an important activity of program development, which is applied to detect incorrect computations or runtime failures of software. The costs of debugging are strongly related to the complexity and the scale of the investigated programs. Both characteristics are especially cumbersome for large-scale parallel programs with long runtimes, which are quite common in computational science and engineering (CSE) applications. A solution is offered by a combination of techniques using the event graph model as a representation of parallel program behaviour. With process isolation, a subset of the original number of processes can be investigated, while the absent processes are simulated by the debugging system. With checkpointing, an arbitrary temporal section of a program's runtime can be extracted for exhaustive analysis without the need to restart the program from the beginning. Additional benefits of the event graph are support of equivalent execution of nondeterministic programs, as well as a comprehensible visualisation as a space-time diagram. (C) 2002 Elsevier Science B.V. All rights reserved.
Keyword:
error detection
space-time diagrams
parallel program
debugging
event graph
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
F
IF:
6.1
论文数:
6.9K
被引数:
2.3W
机构
暂无机构信息
引用论文
Simulation of CO2 Fluxes in European Forest Ecosystems with the Coupled Soil-Vegetation Process Model “LandscapeDNDC”
Forests
IF0

