返回
Autotuning GEMM Kernels for the Fermi GPU
DOI:10.1109/TPDS.2011.311.png)
摘要
En 中文
In recent years, the use of graphics chips has been recognized as a viable way of accelerating scientific and engineering applications, even more so since the introduction of the Fermi architecture by NVIDIA, with features essential to numerical computing, such as fast double precision arithmetic and memory protected with error correction codes. Being the crucial component of numerical software packages, such as LAPACK and ScaLAPACK, the general dense matrix multiplication routine is one of the more important workloads to be implemented on these devices. This paper presents a methodology for producing matrix multiplication kernels tuned for a specific architecture, through a canonical process of heuristic autotuning, based on generation of multiple code variants and selecting the fastest ones through benchmarking. The key contribution of this work is in the method for generating the search space; specifically, pruning it to a manageable size. Performance numbers match or exceed other available implementations.
Keyword:
Graphics processing unit
matrix multiplication
code generation
automatic tuning
GEMM
BLAS
CUDA
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
Genome-Wide Association Study and Subsequent Exclusion of ATCAY as a Candidate Gene Involved in Equine Neuroaxonal Dystrophy Using Two Animal Models
Genes
IF0

