Return
Better than Optimal: Improving Adaptive Stochastic Quantization Using Shared Randomness
DOI:10.1145/3771564.png)
Abstract
En 中文
Quantization is used to convert higher-precision data into a lower-precision representation, providing a fundamental technique to reduce memory usage, computational complexity, and communication costs in machine learning tasks. Recently, Adaptive Stochastic Quantization (ASQ) methods, which provide unbiased optimized quantized values for a specific input (i.e., when the input distribution is not known to the dequantizer), have been shown to be efficient for compressing large tensors on the fly. In this paper, we introduce the Adaptive Unbiased Quantization (AUQ) problem that generalizes ASQ to a setting where shared randomness known to both the quantizer and the dequantizer is allowed. We prove an asymptotic gap between the achievable accuracy of the problems, showing that AUQ can be quadratically more precise. We then introduce Simba, an AUQ technique that improves the accuracy of optimal ASQ methods. Intuitively, Simba uses multiple sets of quantization values, and the set used to quantize a specific value is determined using the shared randomness. Interestingly, we show that Simba improves both the speed and accuracy of optimal ASQ methods by leveraging an approximate ASQ solution as a starting point from which Simba improves. We measure Simba to be up to approximate to 24x more accurate for popular distributions than the state-of-the-art optimal ASQ method, QUIVER, while also being up to two orders of magnitude faster. Finally, we demonstrate that Simba can be used to improve the accuracy of QUIC-FL, a state-of-the-art Distributed Mean Estimation technique, by replacing its solver-based lookup tables with tables computed by Simba.
Keywords:
Compression
Quantization
Shared Randomness
Machine Learning
Stochastic Quantization
Unbiasedness
Journal
P
IF:
2.7
Papers:
45
Citations:
1.0K

