Return
DASA: Distribution-Aware Sparse Attention for Accelerating Diffusion Transformer
DOI:10.1109/TCAD.2025.3627864.png)
Abstract
En 中文
Diffusion transformers (DiTs) have demonstrated remarkable success in text-to-video generation. However, the self-attention mechanism in DiTs imposes significant computational and memory burdens, particularly when handling long patch sequences like high-resolution or long-time videos. While sparse attention shows promise in reducing self-attention costs, existing approaches struggle to deliver performance gains due to the unique challenges in DiTs, i.e., varied sparse patterns across layers and timesteps, and the cumulative nature of inference errors over timesteps. In this article, we propose DASA, an algorithm-hardware co-design that effectively addresses these challenges of attention sparsification in DiTs. Specifically, leveraging the insight that the generation quality is primarily influenced by overall distribution drift rather than changes in specific values, we introduce a novel distribution-aware filtering (DAF) mechanism for sparsification. To further accelerate the process, we design a specialized filtering unit that enables fast candidate selection based on the proposed DAF mechanism. Experimental results show that DASA achieves $2.52\times $ speed up compared to A100 GPU, and up to $1.22\times $ speedup over state-of-the-art accelerators for self-attention computation.
Keywords:
Algorithm-hardware co-design
diffusion transformer (DiT)
distribution-aware filtering (DAF)
sparse attention
Journal
I
IF:
2.9
Papers:
586
Citations:
9.6K

