arrow
返回

Coded Deep Learning: Framework and Algorithm

delete2025-11-01
delete0
PRE
AI
E
En‐hui Yang *
S
Shayan Mohajer Hamidi
DOI:10.1109/TIT.2025.3610095delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The success of deep learning (DL) is often achieved at the expense of large model sizes and high computational complexity during both training and post-training inferences, making it difficult to train and run large models in a resource-limited environment. To alleviate these issues, this paper introduces a new framework dubbed coded deep learning (CDL), which integrates information-theoretic coding concepts into the inner workings of DL, aiming to substantially compress model weights and activations, reduce computational complexity at both training and post-training inference stages, and enable efficient model/data parallelism. Specifically, within CDL, 1) we first propose a novel probabilistic method for quantizing both model weights and activations, and its soft differentiable variant which offers an analytic formula for gradient calculation during training; 2) both the forward and backward passes during training are executed over quantized weights and activations, which eliminates a majority of floating-point operations and reduces the training computation complexity; 3) during training, both weights and activations are entropy constrained so that they are compressible in an information-theoretic sense at any stage of training, which in turn reduces communication costs in cases where model/data parallelism is adopted; and 4) the trained model in CDL is by default in a quantized format with compressible quantized weights, reducing post-training inference complexity and model storage complexity. Additionally, a variant of CDL, namely relaxed CDL (R-CDL), is presented to further improve the trade-off between validation accuracy and compression at the disadvantage of full precision operation involved in forward and backward passes during training with other advantageous features of CDL intact. Extensive empirical results show that CDL and R-CDL outperform the state-of-the-art algorithms in DNN compression in the literature.
Keyword:
Training
Computational modeling
Quantization (signal)
Entropy
Parallel processing
Information theory
Data models
Probabilistic logic
Analytical models
Deep learning
entropy
Huffman coding
model and data parallelism
quantization-aware training

期刊

I
IEEE Transactions on Information Theory
IF:
2.9
论文数:
317
被引数:
0

机构

U
university of waterloo
学者数:
2.6K
论文数: 1.4K
被引数: 1
引用论文

引用论文

Mixed-Precision Neural Network Quantization via Learned Layer-Wise Importance
err2022-11-03
err0
PREAI
errChen Tang; Kai Ouyang; Zhi Wang; Yifei Zhu; Wen Ji; Yaowei Wang; Wenwu Zhu
err分享
err收藏
Deep Network Quantization via Error Compensation
err2022-09-01
err10
PREAI
errPeng, Hanyu; Wu, Jiaxiang; Zhang, Zhiwei; Chen, Shifeng; Zhang, Hai-Tao
err分享
err收藏
Universal Deep Neural Network Compression
err2020-05-01
err46
errOAAI
errChoi, Yoojin; El-Khamy, Mostafa; Lee, Jungwon
err分享
err收藏
学者 查看更多内容