arrow
Return

Model Compression via Position-Based Scaled Gradient

delete2022-01-01
delete0
delete
OA
AI
J
Jangho Kim
K
KiYoon Yoo
N
Nojun Kwak *
DOI:10.1109/ACCESS.2022.3231455delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We propose the position-based scaled gradient (PSG) that scales the gradient depending on the position of a weight vector to make it more compression-friendly. First, we theoretically show that applying PSG to the standard gradient descent (GD), which is called PSGD, is equivalent to the GD in the warped weight space, a space made by warping the original weight space via an appropriately designed invertible function. Second, we empirically show that PSG acting as a regularizer to the weight vectors is favorable for model compression domains such as quantization, pruning, and knowledge distillation. PSG reduces the gap between the weight distributions of a full-precision model and its compressed counterpart. This enables the versatile deployment of a model either as an uncompressed mode or as a compressed mode depending on the availability of resources. The experimental results on CIFAR-10/100 and ImageNet datasets show the effectiveness of the proposed PSG in model compression including an iterative pruning method and the knowledge distillation.
Keywords:
Convolutional neural networks
model compression
scaled gradient method
regularization

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

K
kookmin university
Scholars:
3.0K
Papers: 3.3K
Citations: 2
S
seoul national university (snu)
Scholars:
7.2W
Papers: 6.6W
Citations: 86