Return
CaPTQ: Calibration Data Selection for Visual Services Based on Post-Training Quantization
DOI:10.1109/tsc.2026.3689186.png)
Abstract
En 中文
Training high-performance neural network models on the cloud, quantizing them to low-bit, and deploying those low-bit models on ubiquitous Internet of Things (IoT) devices to provide high-quality visual services with low resource requirements and fast response has become mainstream. However, due to the lack of detailed analysis of calibration data, most existing related methods have two limitations: (i) Low-bit models have low performance. (ii) The need for a large amount of calibration data results in a significant consumption of resources. To this end, this paper first proposes a calibration data selection method, CaPTQ, which selects calibration data with smooth pixel distributions to reduce rounding loss and reconstruction loss in the quantization process, thereby improving the performance of the quantized model. Then, this paper proposes a sampling strategy for calibration data without replacement to ensure that the calibration data sampled each time is not repeated, effectively reducing the amount of calibration data used while ensuring the accuracy of the quantized model. The key idea of this paper is to fully analyze the impact of the distribution and quantity of calibration data on the performance of the post-training quantization (PTQ) method and propose a calibration data selection strategy to reduce the performance loss of the quantized model and the amount of calibration data used. Finally, experimental results on the ImageNet-1K and MS COCO datasets for image classification and object detection services demonstrate that the proposed method outperforms other state-of-the-art methods. Specifically, on ImageNet-1K, when the MobileNetV2 model is quantized to W3A3, the accuracy of CaPTQ is approximately 3.7% higher than that of BRECQ. When using the same amount of calibration data, CaPTQ achieves approximately 4.6% higher accuracy than QDrop when quantizing ResNet-101 to W2A2.
Keywords:
Service computing
neural network model quantization
device-cloud collaboration
IoT devices
Journal
IF:
5.8
Papers:
2.1K
Citations:
6.5K

