arrow
Return

Conditional Temporal Variational AutoEncoder for Action Video Prediction

delete2023-06-18
delete1
PRE
AI
X
Xiaogang Xu *
Y
Yi Wang
L
Liwei Wang
B
Bei Yu
J
Jiaya Jia
DOI:10.1007/s11263-023-01832-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To synthesize a realistic action sequence based on a single human image, it is crucial to model both motion patterns and diversity in the action video. This paper proposes an Action Conditional Temporal Variational AutoEncoder (ACT-VAE) to improve motion prediction accuracy and capture movement diversity. ACT-VAE predicts pose sequences for an action clip from a single input image. It is implemented as a deep generative model that maintains temporal coherence according to the action category with a novel temporal modeling on latent space. Further, ACT-VAE is a general action sequence prediction framework. When connected with a plug-and-play Pose-to-Image network, ACT-VAE can synthesize image sequences. Extensive experiments bear out our approach can predict accurate pose and synthesize realistic image sequences, surpassing state-of-the-art approaches. Compared to existing methods, ACT-VAE improves model accuracy and preserves diversity.
Keywords:
Variational AutoEncoder
Action modeling
Temporal coherence
Adversarial learning

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

S
Shanghai Artificial Intelligence Laboratory
Scholars:
470
Papers: 258
Citations: 765
C
Chinese University of Hong Kong
Scholars:
3.4W
Papers: 3.2W
Citations: 5.6W
Z
Zhejiang Laboratory
Scholars:
1.8K
Papers: 1.7K
Citations: 0
researcher View more organizations