arrow
Return

LogFlex: Flexible-Bit Log Arithmetic Accelerator for Language Models on Edge

delete2026-02-27
delete0
PRE
AI
F
Faraz Tahmasebi
G
Gunjae Koo
H
Hyoukjun Kwon
DOI:10.1109/MM.2026.3669122delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deploying language models on resource-constrained mobile/wearable devices while maintaining output quality is challenging. To address such challenges, many floating-point (FP) and integer (INT) quantization methods have been explored. FP arithmetic provides high quality at the cost of heavy area and energy costs, while INT-quantized models deliver superior efficiency at the cost of accuracy/perplexity loss. In addition to the design choices between FP and INT, we explore an alternative option based on a logarithmic number system (LNS), which delivers FP-like accuracy/perplexity at an efficiency close to INT. In addition to applying low-precision (8-bit) LNS, we adaptively assign bits for the INT and the fraction depending on data distribution, which enables near-FP16 accuracy/perplexity. We also co-design the LNS arithmetic and accelerator architecture, which leads to 33% less energy than the FP8 (E4M3) accelerator with similar area as an INT8 accelerator, while delivering 30% lower perplexity compared to FP8 (E4M3).
Keywords:
Quantization (signal)
Hardware
Artificial intelligence
Arithmetic
Costs
Tensors
Dynamic range
Computational modeling
Accuracy
Partitioning algorithms

Journal

IEEE Micro cover
IEEE Micro
IF:
2.9
Papers:
125
Citations:
2.7K

Organization

K
korea university
Scholars:
4.7K
Papers: 2.1K
Citations: 1
U
University of California
Scholars:
8.7K
Papers: 3.3K
Citations: 8.3W