arrow
Return

Safety and reliability driven task allocation in distributed systems

delete1999-03-01
delete97
PRE
AI
S
Srikanth Srinivasan *
N
Niraj K. Jha
DOI:10.1109/71.755824delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Distributed computer systems are increasingly being employed for critical applications, such as aircraft control, industrial process control, and banking systems. Maximizing performance has been the conventional objective in the allocation of tasks for such systems. Inherently, distributed systems are more complex than centralized systems. The added complexity could increase the potential for system failures. Some work has been done in the past in allocating tasks to distributed systems, considering reliability as the objective function to be maximized. Reliability is defined to be the probability that none of the system components fails while processing. This, however, does not give any guarantees as to the behavior of the system when a failure occurs. A failure, not detected immediately, could lead to a catastrophe. Such systems are unsafe. In this paper, we describe a method to determine an allocation that introduces safely into a heterogeneous distributed system and at the same time attempts to maximize its reliability. First, we devise a new heuristic, based on the concept of clustering to allocate tasks for maximizing reliability. We show that for task graphs with precedence constraints, our heuristic performs better than previously proposed heuristics. Next, by applying the concept of task-based fault tolerance, which we have previously proposed, we add extra assertion tasks to the system to make it safe. We present a new heuristic that does this in such a way that the decrease in reliability for the added safety is minimized. For the purpose of allocating the extra tasks, this heuristic performs as well as previously known methods and runs an order of magnitude faster. We present a number of simulation results to prove the efficacy of our scheme.
Keywords:
allocation
clustering
distributed systems
reliability
safety
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

No organization information available
Cited Papers

Cited Papers

Modification of cellulose model surfaces by cationic polymer latexes prepared by RAFT-mediated surfactant-free emulsion polymerization
err2014-06-30
err0
errOAAI
errLinn Carlsson; Andreas Fall; Isabelle Chaduc; Lars Wågberg; Bernadette Charleux; Eva Malmström; Franck D'Agosto; Muriel Lansalot; Anna Carlmark
errShare
errSave
A nonconcerted cycloaddition of fused 2-vinylthiophenes with dimethyl acetylenedicarboxylate
err2004-03-01
err0
PREAI
errAleš Machara; Milan Kurfürst; Václav Kozmı́k; Hana Petřı́čková; Hana Dvořáková; Jiřı́ Svoboda
errShare
errSave
Perceptions of Success in Bariatric Surgery: a Nationwide Survey Among Medical Professionals
err2017-07-10
err0
PREAI
errShiri Sherf-Dagan; Lihi Schechter; Rita Lapidus; Nasser Sakran; David Goitein; Asnat Raziel
errShare
errSave
Effectiveness of Gastric Bypass Versus Gastric Sleeve for Cardiovascular Disease: Protocol and Baseline Results for a Comparative Effectiveness Study
err2020-04-06
err0
errOAAI
errKaren J Coleman; Heidi Fischer; David E Arterburn; Douglas Barthold; Lee J Barton; Anirban Basu; Anita Courcoulas; Cecelia L Crawford; Peter Fedorka; Benjamin Kim; Edward Mun; Sameer Murali; Kristi Reynolds; Kangho Suh; Rong Wei; Tae K Yoon; Robert Zane
errShare
errSave
errShare
errSave
errShare
errSave
researcher View more