arrow
Return

PBBQ: Plug-In Balanced Binary Quantization for LLMs

delete2026-02-13
delete0
delete
OA
AI
Z
Zhangming Li
W
Weifan Guan
Z
Zhengwei Chang
L
Linghao Zhang
胡庆浩 cover
胡庆浩 (Qinghao Hu) *
DOI:10.3390/electronics15040819delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In recent years, the expansion of large-model parameters has substantially increased storage and inference overhead. Consequently, post-training quantization has become a key technique for reducing model size and inference-time energy consumption. However, we observe that, under extremely low bit-width settings, mainstream error-compensation-based algorithms tend to overfit the calibration data. To mitigate this issue, we propose Plug-in Balanced Binary Quantization for LLMs (PBBQ), which reduces the excessive emphasis on subsequent channels via block-wise dropout and layer-wise reordering. PBBQ can be integrated into GPTQ-style frameworks and ultra-low-bit methods such as BiLLM and ARB-LLM. Experimental results show that PBBQ significantly improves the performance of multiple error-compensation quantization algorithms. When combined with the state-of-the-art methods BiLLM and ARB-LLM, the perplexity (ppl) on WikiText-2 is reduced by 21.46% (from 32.48 to 25.51) and 22.02% (from 16.44 to 12.82), respectively.
Keywords:
large model compression
low-bit quantization
post-training quantization
error compensation quantization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Electronics cover
Electronics
IF:
2.6
Papers:
9.6K
Citations:
4.7W

Organization

S
State Grid Sichuan Electric Power Company
Scholars:
59
Papers: 31
Citations: 0
K
key laboratory of sichuan province
Scholars:
41
Papers: 13
Citations: 0
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704
researcher View more organizations