arrow
Return

Reducing LLM Inference Memory Bandwidth via Frequent Exponent Value Encoding

delete2026-01-01
delete0
PRE
AI
M
Maxwell Michalec *
S
Swamit Tannu
G
Gurindar S. Sohi
DOI:10.1109/LCA.2026.3671166delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We consider the issue of the memory bandwidth required for the transfer of weights of an LLM between the memory and the processor (CPU or GPU). Observing that a few exponent values dominate for FP8 and BF16 floating-point formats, we propose an alternate (compressed) representation and storage of the exponents in memory, and a design for reconstructing the full, exact weight values in the processor. The proposed approach reduces the memory bandwidth needed by up to 10% and 30% for LLMs using weight in FP8 and BF16 formats, respectively, without changing any of the computation carried out by the original LLM inference software.
Keywords:
Tensors
Random access memory
Encoding
Decoding
Codes
Software
Computational modeling
Metadata
Entropy
Bandwidth
Artificial intelligence
data compaction and compression
language models

Journal

I
IEEE Computer Architecture Letters
IF:
1.4
Papers:
42
Citations:
781

Organization

University of Wisconsin System cover
University of Wisconsin System
Scholars:
6.7W
Papers: 5.8W
Citations: 382