arrow
返回

LSTD: Long Short-Term Temporal Diffusion for Video Generation

delete2026-01-12
delete0
PRE
AI
H
Haoyu Zhao
J
Jiaxi Gu
S
Shicong Wang
T
Tianyi Lu
X
Xing Zhang
Z
Zuxuan Wu
H
Hang Xu
Y
Yu–Gang Jiang
DOI:10.1109/TMM.2026.3651052delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Recently, text-driven video generation has achieved tremendous progress. However, existing methods neglect the contexts of long short-term frames in the video, thereby compromising temporal consistency. They also encounter challenges of heavy memory costs due to the use of the standard temporal attention mechanism and misalignment between training videos and captions. Additionally, previous approaches for long video generation are flawed because they are hard to ensure content diversity and consistency. To alleviate these issues, we propose a novel Long Short-term Temporal Diffusion (LSTD) model to generate videos with superior temporal consistency. We introduce two novel temporal modules, i.e., the Short-term Temporal Convolution and the Long-term Temporal Attention. The former can learn short-term features with a shallow structure, and the latter concentrates on long-term information of complex motion with a new memory-efficient attention mechanism. The combination of the two modules can ensure the temporal consistency of the generated videos. Furthermore, a novel inference method for long video generation is also proposed, which can iteratively generate hundreds of video frames. Experimental results on UCF-101, MSR-VTT, and two long video benchmarks prove that our method achieves superior zero-shot inference performance even when the size of the training data is reduced by 26.5 times.
Keyword:
Long short-term temporal diffusion
long video generation
text-to-video
video generation

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

T
tencent
学者数:
63
论文数: 28
被引数: 0
F
fudan university
学者数:
11.8W
论文数: 7.7W
被引数: 121
H
huawei
学者数:
43
论文数: 15
被引数: 0
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
err2023-10-01
err0
errOAAI
errSongwei Ge; Seungjun Nah; Guilin Liu; Tyler Poon; Andrew Tao; Bryan Catanzaro; David Jacobs; Jia-Bin Huang; Ming-Yu Liu; Yogesh Balaji
err分享
err收藏
Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding
err2022-12-01
err0
errOAAI
errMathew Monfort; Bowen Pan; Kandan Ramakrishnan; Alex Andonian; Barry A. McNamara; Alex Lascelles; Quanfu Fan; Dan Gutfreund; Rogerio Schmidt Feris; Aude Oliva
err分享
err收藏
Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
err2022-10-24
err0
PREAI
errSongwei Ge; Thomas Hayes; Harry Yang; Xi Yin; Guan Pang; David Jacobs; Jia-Bin Huang; Devi Parikh
err分享
err收藏
TA2V: Text-Audio Guided Video Generation
err2024-01-01
err1
PREAI
errZhao, Minglu; Wang, Wenmin; Chen, Tongbao; Zhang, Rui; Li, Ruochen
err分享
err收藏
Sounding Video Generator: A Unified Framework for Text-Guided Sounding Video Generation
err2024-01-01
err2
errOAAI
errLiu, Jiawei; Wang, Weining; Chen, Sihan; Zhu, Xinxin; Liu, Jing
err分享
err收藏
LaVie: High-Quality Video Generation with Cascaded Latent Diffusion ModelsLaVie: 具有级联潜在扩散模型的高质量视频生成
err2024-12-23
err0
PREAI
errWang, Yaohui; Chen, Xinyuan; Ma, Xin; Zhou, Shangchen; Huang, Ziqi; Wang, Yi; Yang, Ceyuan; He, Yinan; Yu, Jiashuo; Yang, Peiqing; Guo, Yuwei; Wu, Tianxing; Si, Chenyang; Jiang, Yuming; Chen, Cunjian; Loy, Chen Change; Dai, Bo; Lin, Dahua; Qiao, Yu; Liu, Ziwei
err分享
err收藏
学者 查看更多内容