arrow
Return

Scripted Video Generation With a Bottom-Up Generative Adversarial Network

delete2020-01-01
delete14
PRE
AI
Q
Qi Chen
Q
Qi Wu
陈健 cover
陈健 (Jian Chen)
吴庆耀 (Qingyao Wu)
A
Anton van den Hengel
谭明奎 cover
谭明奎 (Mingkui Tan) *
DOI:10.1109/TIP.2020.3003227delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Generating videos given a text description (such as a script) is non-trivial due to the intrinsic complexity of image frames and the structure of videos. Although Generative Adversarial Networks (GANs) have been successfully applied to generate images conditioned on a natural language description, it is still very challenging to generate realistic videos in which the frames are required to follow both spatial and temporal coherence. In this paper, we propose a novel Bottom-up GAN (BoGAN) method for generating videos given a text description. To ensure the coherence of the generated frames and also make the whole video match the language descriptions semantically, we design a bottom-up optimisation mechanism to train BoGAN. Specifically, we devise a region-level loss via attention mechanism to preserve the local semantic alignment and draw details in different sub-regions of video conditioned on words which are most relevant to them. Moreover, to guarantee the matching between text and frame, we introduce a frame-level discriminator, which can also maintain the fidelity of each frame and the coherence across frames. Last, to ensure the global semantic alignment between whole video and given text, we apply a video-level discriminator. We evaluate the effectiveness of the proposed BoGAN on two synthetic datasets (i.e., SBMG and TBMG) and two real-world datasets (i.e., MSVD and KTH).
Keywords:
Generative adversarial networks
video generation
semantic alignment
temporal coherence
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

U
University of Adelaide
Scholars:
2.3W
Papers: 2.4W
Citations: 4.2W
S
south china university of technology
Scholars:
6.7W
Papers: 5.1W
Citations: 85