Return
DSLA: An Energy-Efficient Dual-Sparsity LLM Accelerator With HiMix-BFP
Z
S
Y
K
韩
DOI:10.1109/tcsi.2026.3692866.png)
Abstract
En 中文
Large language models (LLMs) have achieved great success in areas such as language understanding and text generation. However, their massive parameters incur high computational, storage, and energy costs, making deployment on resource- and power-constrained edge devices particularly challenging. Block Floating Point (BFP) reduces storage and computational overhead by grouping data into blocks and aligning them to the maximum exponent within each block, converting them into low-bit fixed-point numbers. Bidirectional Block Floating Point (BBFP) extends BFP by aligning data within a block to two different exponents, reducing quantization error for small values. However, at ultra-low bit widths, both methods suffer from severe accuracy degradation due to their sensitivity to outliers, which limits their ability to achieve aggressive energy savings. In addition, the exponent alignment procedure in these schemes inherently introduces bit-sliced sparsity, an opportunity that remains largely unexplored for further improving energy efficiency on edge accelerators. To address these challenges, we propose an energy-efficient dual-sparsity LLM accelerator (DSLA) that supports the HiMix-BFP data format. HiMix-BFP improves low-bit accuracy by preserving extra mantissa bits for the maximum value and adaptively selecting BFP or BBFP per block. The DSLA architecture efficiently exploits both value and bit-sliced sparsity across different bit widths to further enhance energy efficiency by employing low-bit computational units, load-balancing mechanisms, and a hierarchical bit-accumulation array. Experimental results demonstrate that HiMix-BFP reduces perplexity by up to 76%, while DSLA achieves up to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1.39\times $ </tex-math></inline-formula> higher throughput and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1.59\times $ </tex-math></inline-formula> greater energy efficiency compared to SOTA accelerators.
Keywords:
Large language models
block floating point
bit-sparse accelerator
Journal
IF:
5.2
Papers:
9.7K
Citations:
2.2W
