Return
Reducing LLM Inference Memory Bandwidth via Frequent Exponent Value Encoding
DOI:10.1109/LCA.2026.3671166.png)
Abstract
En 中文
We consider the issue of the memory bandwidth required for the transfer of weights of an LLM between the memory and the processor (CPU or GPU). Observing that a few exponent values dominate for FP8 and BF16 floating-point formats, we propose an alternate (compressed) representation and storage of the exponents in memory, and a design for reconstructing the full, exact weight values in the processor. The proposed approach reduces the memory bandwidth needed by up to 10% and 30% for LLMs using weight in FP8 and BF16 formats, respectively, without changing any of the computation carried out by the original LLM inference software.
Keywords:
Tensors
Random access memory
Encoding
Decoding
Codes
Software
Computational modeling
Metadata
Entropy
Bandwidth
Artificial intelligence
data compaction and compression
language models
Journal
I
IF:
1.4
Papers:
42
Citations:
781


