arrow
Return

DiffAccel: Accelerating Diffusion Models Through Adaptive Feature Optimization and Dynamic Hardware Adaptation

delete2025-09-08
delete0
PRE
AI
E
Enhao Tang
W
W. Ma
Y
Yudan Jiang
S
Sheng Xu
Z
Zhongfeng Wang
J
Jun Lin
DOI:10.1109/TVLSI.2025.3604419delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Diffusion models (DMs) have emerged as a revolutionary technology in AI-generated content. Despite their superior quality, DMs face significant deployment challenges due to two key limitations: 1) the substantial computational cost inherent in their iterative denoising process and 2) the complexities of optimizing the execution of their multiscale U-Net architecture, which exhibits diverse computational patterns and varying memory requirements across layers. Existing optimization approaches face limitations on both algorithmic and hardware fronts. At the algorithm level, current methods struggle with inefficient feature enhancement through uniform channel processing and inflexible feature reuse with static caching strategies. On the hardware side, current solutions lack the flexibility and efficiency to adapt to the varied computational and memory demands within the U-Net. To address these challenges, we propose DiffAccel, an algorithm–hardware co-optimized accelerator designed to accelerate the inference of various DMs. First, two key algorithmic innovations are introduced: selective feature enhancement targeting only high-dimensional channels and an adaptive feature caching mechanism based on identified evolution patterns. Second, at the hardware level, a reconfigurable array (RA) is proposed to enable efficient execution across diverse U-Net layers by switching its dataflow between output-stationary (OS) and weight-stationary (WS) modes based on the specific computational patterns of each layer. Moreover, a data-volume-aware memory management unit (MMU) is implemented, employing adaptive input allocation and hierarchical output storage to efficiently manage the disparate layer memory footprints of the U-Net, thereby improving on-chip memory efficiency and reducing waste. Third, fine-grained pipeline fusion is employed for key U-Net operation sequences to minimize intermediate data movement and significantly reduce latency. Specifically, row-first processing is applied for linear–Softmax operations, and channel-priority output is utilized for convolution–groupnorm (GNorm) sequences. At the algorithm level, DiffAccel achieves $2.93\times $ speedup and reduces computational complexity (from 105 to 69 TFLOPs) compared to the baseline stable diffusion (SD) v2.0. DiffAccel is implemented with a 28-nm technology. In terms of hardware efficiency, based on post place-and-route (P&R) implementation results, DiffAccel achieves $5.25\times $ – $5.79\times $ higher throughput and maintains comparable area efficiency compared to state-of-the-art (SOTA) accelerators.
Keywords:
Diffusion models (DMs)
feature optimization
hardware adaptation

Journal

I
IEEE Transactions on Very Large Scale Integration (VLSI) Systems
IF:
3.1
Papers:
440
Citations:
7.3K

Organization

N
nanjing university
Scholars:
7.7W
Papers: 5.6W
Citations: 87
S
Southeast University
Scholars:
2.0W
Papers: 8.3K
Citations: 480