返回
A Vector Systolic Accelerator for Multi-Precision Floating-Point High-Performance Computing
DOI:10.1109/TCSII.2022.3183007.png)
摘要
En 中文
There is an emerging need to design multi-precision floating-point (FP) accelerators for high-performance-computing (HPC) applications. The commonly-used methods are based on high-precision-split (HPS) and low-precision-combination (LPC) structures, which suffer from low hardware utilization ratio and various multiple clock-cycle processing periods. In this brief, a new multi-precision FP processing element (PE) is developed with proposed bit-partitioning method. Minimized redundant bits and operands are achieved. The proposed PE supports 16x half-precision (FP16), 4x single-precision (FP32) and 1x double-precision (FP64) operations with 100% multiplication hardware utilization ratio. Besides, vector systolic structure is designed for PE array to increase the system-level throughput and energy efficiency. The proposed design is realized in a 28-nm process with 1.351-GHz clock frequency. Compared with the existing multi-precision FP methods, the proposed work exhibits the best energy-efficiency performance of 1193 GFLOPS/W at FP16, 317 GFLOPS/W at FP32 and 77.3 GFLOPS/W at FP64 with at least 22.3%, 30% and 3.3% improvement, respectively.
Keyword:
Multi-precision
floating-point
PE
MAC
vector
systolic
HPC
accelerator
期刊
I
IF:
4.9
论文数:
8.8K
被引数:
2.5W
机构
暂无机构信息
引用论文
Review and Benchmarking of Precision-Scalable Multiply-Accumulate Unit Architectures for Embedded Neural-Network Processing用于嵌入式神经网络处理的精度可扩展乘法累积单元体系结构的审查和基准测试

