Return
Precision boundary modeling for area-efficient Block Floating Point accumulation
DOI:10.1016/j.sysarc.2026.103704.png)
Abstract
En 中文
Block Floating Point (BFP) is extensively employed for the low-precision quantization of deep network weights and activations to attain advantages in both hardware efficiency and performance. Nevertheless, when the precision of weights and activations is diminished to below 8 bits, the required high-precision floating-point accumulation becomes a dominant hardware bottleneck in the BFP processing element (PE). To address this challenge, we introduce a framework based on the Frobenius norm Retention Ratio (FnRR) to explore the precision boundaries for BFP accumulation, and extend it to a hierarchical chunk-based accumulation scheme. Comprehensive experiments across representative CNN and LLM models demonstrate that our predicted precision boundaries maintain performance closely matching FP32 baselines, while further precision reduction leads to substantial accuracy degradation, validating the effectiveness of our boundary determination. Guided by this analysis, we present a corresponding hardware for BFP computation. This design achieves 13.7%–25.2% improvements in area and power efficiency compared with FP32 accumulation under identical quantization settings, and delivers up to 10.3× area and 11.0× power reductions relative to conventional BFP implementations.
Journal
IF:
4.1
Papers:
3.0K
Citations:
4.2K

