arrow
Return

Memory-Efficient Batch Normalization by One-Pass Computation for On-Device Training

delete2024-06-01
delete0
PRE
AI
H
H. Y. Dai
H
Hang Wang
张续冲 (Xuchong Zhang)
孙宏滨 (Hongbin Sun) *
DOI:10.1109/TCSII.2024.3354738delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Batch normalization (BN) has become ubiquitous in modern deep learning architectures because of its remarkable improvement in deep neural network (DNN) training performance. However, the two-pass computation of statistical estimation and element-wise normalization in BN training requires two accesses to the input data, resulting in a huge increase in off-chip memory traffic during DNN training. In this brief, we propose a novel accelerator, named one-pass normalizer (OPN) to achieve memory-efficient BN for on-device training. Specifically, in terms of dataflow, we propose one-pass computation based on sampling-based range normalization and sparse data recovery techniques to reduce BN off-chip memory access. Regarding the OPN circuit, we propose channel-wise constant extraction to achieve a compact design. Experimental results show that the one-pass computation reduces off-chip memory access of BN by 2.0 similar to 3.8x compared with the previous state-of-the-art designs while maintaining training performance. Moreover, the channel-wise constant extraction saves the gate count and power consumption of OPN by 56% and 73%, respectively.
Keywords:
Training
Systolic arrays
Backpropagation
Artificial neural networks
Micromechanical devices
Feedforward systems
Memory management
Memory-efficient accelerator
deep neural networks
batch normalization
on-device training
one-pass computation

Journal

I
IEEE Transactions on Circuits and Systems and Express Briefs
IF:
4.9
Papers:
8.8K
Citations:
2.5W

Organization

X
xi'an jiaotong university
Scholars:
9.2W
Papers: 6.6W
Citations: 75