返回
LE-GEMM: A lightweight emulation-based GEMM with precision refinement on GPU
DOI:10.1016/j.sysarc.2025.103336.png)
摘要
En 中文
Many special hardware units, such as Matrix Core and Tensor Core, have recently been designed and applied in various scientific computing scenarios. These units support tensor-level computation with different precisions on GPU. Previous studies have proposed methods for computing single-precision GEneral Matrix Multiplication (GEMM) with the half-precision matrix. However, this routine often leads to some loss of accuracy, which limits its application. This paper proposed a Lightweight Emulation-based GEMM (LE-GEMM) on GPU that includes a lightweight emulation algorithm, a thread parallelism analytic model, and an efficient multi-level pipeline implementation to accelerate the computation process without compromising the accuracy requirements. First, we propose a lightweight emulation algorithm that includes a precision transformation process and GEMM emulation calculation to achieve better computational accuracy and performance. Secondly, a thread parallel analytic model is designed to analyze and guide the selection of the optimal tiling scheme based on various computing scenarios and hardware. Thirdly, an efficient multi-level pipeline is implemented, which can maximize instruction-level parallelism and latency hiding. Several comparison experiments were conducted on two commonly used GPU platforms: AMD-platform and NVIDIA-platform. The experimental results show that the proposed method outperforms the previous approaches in terms of computational accuracy and speed.
Keyword:
GEMM
Thread parallelism analytic
Multi-level pipeline
Matrix core/Tensor core
期刊
IF:
4.1
论文数:
3.0K
被引数:
4.2K
机构
引用论文
Investigation of Hot-Carrier-Induced Degradation Mechanisms in p-Type High-Voltage Drain Extended Metal–Oxide–Semiconductor Transistors研究热载流子诱导的p型高压外延漏极金属-氧化物-半导体晶体管中的退化机制

