1
Return

DSLA: An Energy-Efficient Dual-Sparsity LLM Accelerator With HiMix-BFP

delete2026-05-18
delete0
PRE
AI
Z
Zikang Zhou
S
Siyao Dai
Y
Yaqi Chen
K
Kaiqi Chen
韩军 (Jun Han)
DOI:10.1109/tcsi.2026.3692866delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) have achieved great success in areas such as language understanding and text generation. However, their massive parameters incur high computational, storage, and energy costs, making deployment on resource- and power-constrained edge devices particularly challenging. Block Floating Point (BFP) reduces storage and computational overhead by grouping data into blocks and aligning them to the maximum exponent within each block, converting them into low-bit fixed-point numbers. Bidirectional Block Floating Point (BBFP) extends BFP by aligning data within a block to two different exponents, reducing quantization error for small values. However, at ultra-low bit widths, both methods suffer from severe accuracy degradation due to their sensitivity to outliers, which limits their ability to achieve aggressive energy savings. In addition, the exponent alignment procedure in these schemes inherently introduces bit-sliced sparsity, an opportunity that remains largely unexplored for further improving energy efficiency on edge accelerators. To address these challenges, we propose an energy-efficient dual-sparsity LLM accelerator (DSLA) that supports the HiMix-BFP data format. HiMix-BFP improves low-bit accuracy by preserving extra mantissa bits for the maximum value and adaptively selecting BFP or BBFP per block. The DSLA architecture efficiently exploits both value and bit-sliced sparsity across different bit widths to further enhance energy efficiency by employing low-bit computational units, load-balancing mechanisms, and a hierarchical bit-accumulation array. Experimental results demonstrate that HiMix-BFP reduces perplexity by up to 76%, while DSLA achieves up to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1.39\times $ </tex-math></inline-formula> higher throughput and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$1.59\times $ </tex-math></inline-formula> greater energy efficiency compared to SOTA accelerators.
Keywords:
Large language models
block floating point
bit-sparse accelerator

Journal

IEEE Transactions on Circuits and Systems I-Regular Papers cover
IEEE Transactions on Circuits and Systems I-Regular Papers
IF:
5.2
Papers:
9.7K
Citations:
2.2W

Organization

F
fudan university
Scholars:
11.3W
Papers: 7.6W
Citations: 121
Cited Papers

Cited Papers

Citing Papers

Citing Papers