arrow
返回

VecQ: Minimal Loss DNN Model Compression With Vectorized Weight Quantization

delete2021-05-01
delete37
delete
OA
AI
C
Cheng Gong
Y
Yao Chen
Y
Ye Lu *
T
Tao Li *
H
Hao, Cong
D
Deming Chen *
DOI:10.1109/TC.2020.2995593delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Quantization has been proven to be an effective method for reducing the computing and/or storage cost of DNNs. However, the trade-off between the quantization bitwidth and final accuracy is complex and non-convex, which makes it difficult to be optimized directly. Minimizing direct quantization loss (DQL) of the coefficient data is an effective local optimization method, but previous works often neglect the accurate control of the DQL, resulting in a higher loss of the final DNN model accuracy. In this paper, we propose a novel metric, called Vector Loss. Using this new metric, we decompose the minimization of the DQL to two independent optimization processes, which significantly outperform the traditional iterative L2 loss minimization process in terms of effectiveness, quantization loss as well as final DNN accuracy. We also develop a new DNN quantization solution called VecQ, which provides minimal direct quantization loss and achieve higher model accuracy. In order to speed up the proposed quantization process during model training, we accelerate the quantization process with a parameterized probability estimation method and template-based derivation calculation. We evaluate our proposed algorithm on MNIST, CIFAR, ImageNet, IMDB movie review and THUCNews text data sets with numerical DNN models. The results demonstrate that our proposed quantization solution is more accurate and effective than the state-of-the-art approaches yet with more flexible bitwidth support. Moreover, the evaluation of our quantized models on Salient Object Detection (SOD) tasks maintains comparable feature extraction quality with up to 16x weight size reduction.
Keyword:
Quantization (signal)
Training
Euclidean distance
Computational modeling
Data models
Numerical models
DNN compression
DNN quantization
vectorized weight quantization
low bitwidth
vector loss
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Computers 封面图
IEEE Transactions on Computers
IF:
3.8
论文数:
5.3K
被引数:
9.8K

机构

U
University of Illinois Urbana-Champaign
学者数:
2.4W
论文数: 2.0W
被引数: 35
University of Illinois System 封面图
University of Illinois System
学者数:
6.8W
论文数: 6.2W
被引数: 644
N
nankai university
学者数:
4.8W
论文数: 3.3W
被引数: 74
学者 查看更多机构
引用论文

引用论文

Monkeypox vaccination in the global south: Fighting a war without a weapon
err2023-07-01
err0
errOAAI
errIsaac Olushola Ogunkola; Oyinloye Emmanuel Abiodun; Babatunde Ismail Bale; Emmanuel Ebuka Elebesunu; Somtochukwu Blessing Ujam; Innocent Chimaobi Umeh; Mfoniso Tom-James; Shuaibu Saidu Musa; Emery Manirambona; Salvador B. Evardone; Don Eliseo Lucero-Prisno
err分享
err收藏
Abnormality in power system transient stability control of BESS/STATCOM
err2018-01-19
err0
errOAAI
errJun Liu; Can Su; Xu Wang; Wanliang Fang; Shuanbao Niu; Lin Cheng
err分享
err收藏
Kinetics of Phenol Escape from the Insulin R6 Hexamer
err2021-10-14
err0
errOAAI
errAdam Antoszewski; Chatipat Lorpaiboon; John Strahan; Aaron R. Dinner
err分享
err收藏
Latent mean differences in executive function in at-risk preterm children: The delay-deficit dilemma.
err2014-01-01
err0
PREAI
errIda Sue Baron; Brandi A. Weiss; Fern R. Litman; Margot D. Ahronovich; Robin Baker
err分享
err收藏
学者 查看更多内容