arrow
Return

DyDiT++: Diffusion Transformers With Timestep and Spatial Dynamics for Efficient Visual Generation

delete2026-01-15
delete0
PRE
AI
W
Wangbo Zhao
Y
Yizeng Han
J
Jiasheng Tang
K
Kai Wang
H
Hao Luo
Y
Yibing Song
G
Gao Huang
王繁 (Fan Wang)
Y
Yang You
DOI:10.1109/TPAMI.2026.3654201delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Diffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the <i>static</i> inference paradigm, which inevitably introduces redundant computation in certain <i>diffusion timesteps</i> and <i>spatial regions</i>. To overcome this inefficiency, we propose <b>Dy</b>namic <b>Di</b>ffusion <b>T</b>ransformer (DyDiT), an architecture that <i>dynamically</i> adjusts its computation along both <i>timestep</i> and <i>spatial</i> dimensions. Specifically, we introduce a <i>Timestep-wise Dynamic Width</i> (TDW) approach that adapts model width conditioned on the generation timesteps. In addition, we design a <i>Spatial-wise Dynamic Token</i> (SDT) strategy to avoid redundant computation at unnecessary spatial locations. TDW and SDT can be seamlessly integrated into DiT and significantly accelerate the generation process. Building on these designs, we present an extended version, <b>DyDiT++</b>, with improvements in three key aspects. First, it extends the generation mechanism of DyDiT beyond diffusion to flow matching, demonstrating that our method can also accelerate flow-matching-based generation, enhancing its versatility. Furthermore, we enhance DyDiT to tackle more complex visual generation tasks, including video generation and text-to-image generation, thereby broadening its real-world applications. Finally, to address the high cost of full fine-tuning and democratize technology access, we investigate the feasibility of training DyDiT in a parameter-efficient manner and introduce timestep-based dynamic LoRA (TD-LoRA). Extensive experiments on diverse visual generation models, including DiT, SiT, Latte, and FLUX, demonstrate the effectiveness of DyDiT++. Remarkably, with <inline-formula><tex-math notation="LaTeX">$&lt; $</tex-math></inline-formula>3% additional fine-tuning iterations, our approach reduces the FLOPs of DiT-XL by 51%, yielding 1.73× realistic speedup on hardware, and achieves a competitive FID score of 2.07 on ImageNet.
Keywords:
Diffusion transformer (DiT)
dynamic diffusion transformer (DyDiT)
DyDiT++
visual generation

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

A
alibaba group
Scholars:
1.1K
Papers: 789
Citations: 0
T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137
N
national university of singapore
Scholars:
4.6K
Papers: 2.4K
Citations: 1
researcher View more organizations