arrow
Return

A software scheduling solution to avoid corrupted units on GPUs

delete2016-04-01
delete6
PRE
AI
D
David Defour
E
Eric Petit *
DOI:10.1016/j.jpdc.2016.01.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Massively parallel processors provide high computing performance by increasing the number of concurrent execution units. Moreover, the transistor technology evolves to higher density, higher frequency and lower voltage. The combination of these factors increases significantly the probability of hardware failures. In this paper, we present a methodology to locate and mitigate hardware failures of Nvidia GPUs. Results show that intermittent errors can be precisely localized and have a limited impact to a well defined architecture tile. Therefore, we propose, and demonstrate on a software prototype, a rescheduling strategy to quarantine the defective hardware and ensure correct execution. Our approach significantly improves the GPU fault-tolerance capability and GPU's lifespan, at a reasonable overhead. (C) 2016 Elsevier Inc. All rights reserved.
Keywords:
Reliability
GPGPU
Intermittent error
Scheduling
Fault tolerance
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

U
universite perpignan via domitia
Scholars:
1.2K
Papers: 953
Citations: 1
U
Universite Paris Saclay
Scholars:
7.3W
Papers: 5.3W
Citations: 540