返回
Efficient Replication for Fast and Predictable Performance in Distributed Computing
DOI:10.1109/TNET.2021.3062215.png)
摘要
En 中文
Master-worker distributed computing systems use task replication to mitigate the effect of slow workers on job compute time. The master node groups tasks into batches and assigns each batch to one or more workers. We first assume that the batches do not overlap. Using majorization theory, we show that a balanced replication of batches minimizes the average job compute time for a general class of service time distributions. We then show that the balanced assignment of non-overlapping batches achieves a lower average job compute time than the overlapping schemes proposed in the literature. Next, we derive the optimum redundancy level as a function of the task service time distribution. We show that the redundancy level that minimizes the average job compute time may not coincide with the redundancy level that maximizes job compute time predictability. Therefore, there is a trade-off in optimizing the two metrics. By running experiments on Google cluster traces, we observe that redundancy can reduce the job compute time by one order of magnitude. The optimum level of redundancy depends on the distribution of task service time.
Keyword:
Task analysis
Redundancy
Computational modeling
Machine learning
Internet
Computer architecture
Training
Redundancy
replication
distributed systems
distributed computing
latency
coefficient of variations
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
3.6
论文数:
4.4K
被引数:
9.5K
机构
引用论文
Magnetoresistance in asymmetric ferromagnet/superconductor/ferromagnet double tunnel junctions非对称铁磁体/超导体/铁磁体双隧道结中的磁阻

