Return
Communication Efficient Federated Learning With Quantization-Aware Training Design
DOI:10.1109/TMLCN.2025.3635050.png)
Abstract
En 中文
Model quantization is an effective method that can improve communication efficiency in federated learning (FL). The existing FL quantization protocols almost stay at the level of post-training quantization (PTQ), which comes at the cost of large quantization loss, especially in the setting of low-bits quantization. In this work, we propose a FL quantization training strategy to reduce the impact of quantization on model quality. Specifically, we first apply quantization-aware training (QAT) to FL (QAT-FL), which reduces quantization distortion by adding a fake-quantization module to the model so that the model could perceive future quantization during training. The convergence guarantee of the QAT-FL algorithm is established under certain assumptions. On the basis of the QAT-FL algorithm, we extend the discussion of non-uniform quantization and the adaptive algorithm, so that the model can adaptively adjust the parametric distribution and the number of quantization bits to reduce the amount of traffic in training. Experimental results based on MNIST, CIFAR-10 and FEMNIST datasets show that QAT-FL has advantages in terms of training loss and model inference accuracy, and adaptive-bits quantization of QAT-FL also greatly improves communication efficiency.
Keywords:
Quantization (signal)
Adaptation models
Computational modeling
Training
Convergence
Analytical models
Vectors
Servers
Federated learning
Distortion
quantization-aware learning
non-uniform quantization
post-training quantization
adaptive-bits quantization
Journal
I
IF:
0
Papers:
42
Citations:
0

