arrow
Return

Light-DiT: An Importance-Aware Dynamic Compression Framework for Diffusion Transformers

delete2026-01-01
delete0
PRE
AI
顾成 cover
顾成 (Cheng Gu)
G
Gang Li *
X
Xuan Zhang
J
Jiayao Ling
X
Xiaolong Lin
S
Song, Zhuoran
J
Jian Cheng
梁晓峣 (Xiaoyao Liang)
DOI:10.1007/978-3-031-99857-7_25delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Diffusion Transformers (DiTs) demonstrate remarkable generative abilities in AI. However, the iterative denoising process inherent in diffusion incurs substantial computational cost and memory overhead, impeding fast and energy-efficient edge inference. To mitigate the overhead of denoising, we propose a post-training framework that jointly utilizes pruning and quantization for hardware-efficient DiT inference, which is based on the observation that not all denoising blocks within a model are equally important during image generation. We introduce metrics to assess the importance of DiTs' blocks and layers. To achieve importance-aware dynamic compression, we unify mixed-sparsity pruning and mixed-precision quantization based on importance metrics. Experiments show that our approach achieves a 1.41.x inference speedup through pruning with a mixed precision of W3.2A4.9 while incurring minimal accuracy loss. Furthermore, the evaluation of bit-flexible DNN accelerators demonstrates up to 2.78.x performance improvement, and 1.99.x better energy efficiency can be achieved compared to W8A8 quantization without pruning.
Keywords:
Diffusion Transformers
Model Compression
Hardware-Efficient Inference

Journal

E
EURO-PAR 2025: PARALLEL PROCESSING, PT II
IF:
0
Papers:
24
Citations:
0

Organization

S
shanghai jiao tong university
Scholars:
15.5W
Papers: 11.6W
Citations: 159
C
chinese academy of sciences
Scholars:
56.0W
Papers: 44.8W
Citations: 704