arrow
Return

A Fully Quantized Training Accelerator for Diffusion Network With Tensor Type & Noise Strength Aware Precision Scheduling

delete2024-12-01
delete0
PRE
AI
R
Ruoyang Liu
W
Wenxun Wang
C
Chen Tang
W
Weichen Gao
H
Huazhong Yang
Y
Yongpan Liu *
DOI:10.1109/TCSII.2024.3439319delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Fine-grained mixed-precision fully-quantized methods have great potential to accelerate neural network training, but existing methods exhibit large accuracy loss for more complex models such as diffusion networks. This brief introduces a fully-quantized training accelerator for diffusion networks. It features a novel training framework with tensor-type- and noise-strength-aware precision scheduling to optimize bit-width allocation. The processing cluster design enables dynamical switching bit-width mappings for model weights, allows concurrent processing in 4 different bit-widths, and incorporates a gradient square sum collection unit to minimize on-chip memory access. Experimental results show up to 2.4x training speedup and 81% operation bit-width overhead reduction compared to existing designs, with minimal impact on image generation quality.
Keywords:
Fully-quantized training
mixed-precision quantization
diffusion network
neural network training accelerator

Journal

I
IEEE Transactions on Circuits and Systems and Express Briefs
IF:
4.9
Papers:
8.8K
Citations:
2.5W

Organization

T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137