arrow
Return

SD-Acc: Accelerating Stable Diffusion through Phase-Aware Sampling and Hardware Co-Optimizations

delete2026-07-23
delete0
PRE
AI
Z
Zhican Wang
贺
贺光辉 (Guanghui He)
H
Hongxiang Fan
DOI:10.1109/tc.2026.3715379delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The emergence of diffusion models has significantly enhanced the capabilities of generative AI, leading to improved quality, realism, and creativity in image and video generation. Among these, Stable Diffusion (StableDiff) has become one of the most influential diffusion models for text-to-image generation, serving as a critical component in next-generation multi-modal algorithms. Despite their advances, StableDiff imposes substantial computational and memory demands, affecting inference speed and energy efficiency. To address these hardware performance issues, we identify three key challenges: 1) intensive and potentially redundant computation. 2) heterogeneous operators involving both convolutions and attention mechanisms. 3) widely varied weight and activation sizes. Thus, we present SD-Acc, a novel algorithm and hardware co-optimization solution: At the algorithmic level, we observe that high-level features exhibit high similarity in certain phases of the denoising process, indicating the potential of approximate computation. Therefore, we design an adaptive phase-aware sampling framework to reduce the computational and memory requirements. This general framework automatically explores the trade-off between image quality and algorithmic complexity for any given StableDiff model and user requirements. At the hardware level, we introduce an address-centric dataflow to facilitate the efficient execution of heterogeneous operators within a simple systolic array. The bottleneck of nonlinear operations is comprehensively solved by our novel 2-stage streaming computing and reconfigurable vector processing unit. We further enhance hardware efficiency through an adaptive dataflow optimization that incorporates dynamic reuse and fusion techniques tailored for StableDiff, leading to significant memory access reduction. Across multiple StableDiff models, our extensive experiments demonstrate that our algorithm optimization can achieve up to 3 times reduction in computational demands while preserving the same level of image quality across various models. Together with our highly optimized hardware accelerator, our approach can achieve higher speed and energy efficiency over CPU and GPU implementations.
Keywords:
Diffusion models
phase-aware sampling
algorithm–hardware co-design
hardware acceleration
systolic arrays
dataflow optimization

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.4K
Citations:
9.8K

Organization

S
shanghai jiao tong university
Scholars:
15.7W
Papers: 11.7W
Citations: 159
I
imperial college london
Scholars:
9.9K
Papers: 4.4K
Citations: 0
Cited Papers

Cited Papers

No cited papers available