返回
Straggler Mitigation at Scale
DOI:10.1109/TNET.2019.2946464.png)
摘要
En 中文
Runtime performance variability has been a major issue, hindering predictable and scalable performance in modern distributed systems. Executing requests or jobs redundantly over multiple servers have been shown to be effective for mitigating variability, both in theory and practice. Systems that employ redundancy has drawn significant attention, and numerous papers have analyzed the pain and gain of redundancy under various service models and assumptions on the runtime variability. This paper presents a cost (pain) vs. latency (gain) analysis of executing jobs of many tasks by employing replicated or erasure coded redundancy. The tail heaviness of service time variability is decisive on the pain and gain of redundancy and we quantify its effect by deriving expressions for cost and latency. Specifically, we try to answer four questions: 1) How do replicated and coded redundancy compare in the cost vs. latency tradeoff? 2) Can we introduce redundancy after waiting some time and expect it to reduce the cost? 3) Can relaunching the tasks that appear to be straggling after some time help to reduce cost and/or latency? 4) Is it effective to use redundancy and relaunching together? We validate the answers we found for each of these questions via simulations that use empirical distributions extracted from a Google cluster data.
Keyword:
Task analysis
Redundancy
Runtime
Encoding
Computational modeling
Servers
Distributed computing
Coded and replicated redundancy
straggler relaunch
cost vs latency tradeoff in distributed computing
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
3.6
论文数:
4.4K
被引数:
9.5K
机构
引用论文
Combination of NVP-BEZ235 and Enzastaurin on B-Cell Lymphoma Cell Lines: An Effective Therapeutic Strategy
Blood
IF0
Does the selective serotonin reuptake inhibitor (SSRI) fluoxetine modify canine anxiety related behaviour?选择性5-羟色胺再摄取抑制剂(SSRI)氟西汀是否改变犬类焦虑相关行为?

