arrow
Return

CaPTQ: Calibration Data Selection for Visual Services Based on Post-Training Quantization

delete2026-04-30
delete0
PRE
AI
张俊娜 cover
张俊娜 (Junna Zhang)
P
Ping Yang
C
Chuntao Ding
S
Salman Raza
Q
Qing Zhao
P
Peiyan Yuan
S
Shangguang Wang
DOI:10.1109/tsc.2026.3689186delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Training high-performance neural network models on the cloud, quantizing them to low-bit, and deploying those low-bit models on ubiquitous Internet of Things (IoT) devices to provide high-quality visual services with low resource requirements and fast response has become mainstream. However, due to the lack of detailed analysis of calibration data, most existing related methods have two limitations: (i) Low-bit models have low performance. (ii) The need for a large amount of calibration data results in a significant consumption of resources. To this end, this paper first proposes a calibration data selection method, CaPTQ, which selects calibration data with smooth pixel distributions to reduce rounding loss and reconstruction loss in the quantization process, thereby improving the performance of the quantized model. Then, this paper proposes a sampling strategy for calibration data without replacement to ensure that the calibration data sampled each time is not repeated, effectively reducing the amount of calibration data used while ensuring the accuracy of the quantized model. The key idea of this paper is to fully analyze the impact of the distribution and quantity of calibration data on the performance of the post-training quantization (PTQ) method and propose a calibration data selection strategy to reduce the performance loss of the quantized model and the amount of calibration data used. Finally, experimental results on the ImageNet-1K and MS COCO datasets for image classification and object detection services demonstrate that the proposed method outperforms other state-of-the-art methods. Specifically, on ImageNet-1K, when the MobileNetV2 model is quantized to W3A3, the accuracy of CaPTQ is approximately 3.7% higher than that of BRECQ. When using the same amount of calibration data, CaPTQ achieves approximately 4.6% higher accuracy than QDrop when quantizing ResNet-101 to W2A2.
Keywords:
Service computing
neural network model quantization
device-cloud collaboration
IoT devices

Journal

IEEE Transactions on Services Computing cover
IEEE Transactions on Services Computing
IF:
5.8
Papers:
2.1K
Citations:
6.5K

Organization

H
Henan Academy of Agricultural Sciences
Scholars:
1.9K
Papers: 998
Citations: 1.7K
B
beijing university of posts and telecommunications
Scholars:
2.0K
Papers: 762
Citations: 0
N
national textile university faisalabad
Scholars:
6
Papers: 3
Citations: 0
H
henan normal university
Scholars:
1.1W
Papers: 6.1K
Citations: 6
B
beijing normal university
Scholars:
4.8K
Papers: 2.0K
Citations: 0
researcher View more organizations