arrow
Return

CRAT: Enabling Coordinated Register Allocation and Thread-Level Parallelism Optimization for GPUs

delete2018-06-01
delete9
PRE
AI
X
Xiaolong Xie *
Y
Yun Liang
X
Xiuhong Li
Y
Yudong Wu
G
Guangyu Sun
T
Tao Wang
D
Dongrui Fan
DOI:10.1109/TC.2017.2776272delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The key to the high performance on GPUs lies in the massive threading to enable thread switching and hide long latencies. CPUs are equipped with a large register file to enable fast context switch. However, thread throttling techniques that are designed to mitigate cache contention, lead to under-utilization of registers. Register allocation is a significant factor for performance as it not just determines the single-thread performance, but indirectly affects the TLP. In this paper, we propose Coordinated Register Allocation and Thread-level parallelism (CRAT) to explore the optimization space of register allocation and TLP management on GPUs. CRAT employs both compile-time(CRAT-static) and run-time techniques(CRAT-dyn) to exhaust the design space. CRAT-static works statically to explore TLP and register allocation trade-off and CRAT-dyn exploits dynamic register allocation for further improvement. Experiments indicate that CRAT-static achieves an average 1.25X speedup over existing TLP management technique. On four register-limited applications, CRAT-dyn further improves the performance speedup of CRAT-static from 1.51X to 1.70X.
Keywords:
GPGPU
memory hierarchy
compilers
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

P
peking university
Scholars:
11.8W
Papers: 8.7W
Citations: 146
C
chinese academy of sciences
Scholars:
56.3W
Papers: 44.8W
Citations: 704