arrow
Return

Consistent and Controllable Image Animation With Linear Motion Diffusion Transformers

delete2026-02-12
delete0
PRE
AI
X
Xin Ma
Y
Yaohui Wang
G
Genyun Jia
X
Xinyuan Chen
T
Tien‐Tsin Wong
C
Cunjian Chen
DOI:10.1109/TPAMI.2026.3664227delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated animations remain persistent challenges. While text-to-video (T2V) generation has demonstrated impressive performance with diffusion transformer models, the image animation field still largely relies on U-Net-based diffusion models, which lag behind the latest T2V approaches. Moreover, the quadratic complexity of vanilla self-attention mechanisms in Transformers imposes heavy computational demands, making image animation particularly resource-intensive. To address these issues, we propose MiraMo, a framework designed to enhance efficiency, appearance consistency, and motion smoothness in image animation. Specifically, MiraMo introduces three key elements: (1) A foundational text-to-video architecture replacing vanilla self-attention with efficient linear attention to reduce computational overhead while preserving generation quality; (2) A novel motion residual learning paradigm that focuses on modeling motion dynamics rather than directly predicting frames, improving temporal consistency; and (3) A DCT-based noise refinement strategy during inference to suppress sudden motion artifacts, complemented by a dynamics control module to balance motion smoothness and expressiveness. Extensive experiments against state-of-the-art methods validate the superiority of MiraMo in generating consistent, smooth, and controllable animations with accelerated inference speed. Additionally, we demonstrate the versatility of MiraMo through applications in motion transfer and video editing tasks.
Keywords:
Image animation
linear attention
flow matching

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

S
Shanghai Artificial Intelligence Laboratory
Scholars:
458
Papers: 257
Citations: 765
H
Hefei University of Technology
Scholars:
5.5K
Papers: 1.8K
Citations: 2.1W
M
monash university
Scholars:
8.5K
Papers: 3.8K
Citations: 0
researcher View more organizations