返回
Learning Sparse Convolutional Neural Network via Quantization with Low Rank Regularization
DOI:10.1109/ACCESS.2019.2911536.png)
摘要
En 中文
With the refinement of tasks in artificial intelligence, bringing in exponential level increments in computation cost and storage. Therefore, the augment of computation resource for complicated neural networks severely hinders their applications on limited-power devices in recent years. As a result, there is an impending necessity to compress and accelerate the deep networks by special ways. Considering the different peculiarities of weight quantization and sparse regularization, in this paper, we propose a low rank sparse quantization (LRSQ) method to quantize network weights and regularize the corresponding structures at the same time. Our LRSQ can: 1) obtain low-bit quantized networks to reduce memory and computation cost and 2) learn a compact structure from complex convolutional networks for subsequent channel pruning which has significant reduction on FLOPs. In experimental sections, we evaluate the proposed method on several popular models such as VGG-7/16/19 and ResNet-18/34/50, and results show that this method can dramatically reduce parameters and channels of the network with slight inference accuracy loss. Furthermore, we also visualize and analyze the four-dimensional weight tensors, which shows the low rank and group-sparsity structure of it. Finally, we try pruning unimportant channels which are zero-channels in our quantized model, and finding even a little better precision than the standard full-precision network.
Keyword:
Convolutional neural network (CNN)
weight quantization
spectral regularization
sparsity
visualization
channel pruning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
A Hierarchical Cardiac Rhythm Classification Methodology Based on Electrocardiogram Fiducial Points一种基于心电图基准点的分层心律分类方法
GXNOR-Net: Training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework
NEURAL NETWORKS
IF6.3
没有更多内容

