arrow
Return

Video prediction by efficient transformers

delete2023-02-01
delete21
delete
OA
AI
X
Xi Ye *
G
Guillaume-Alexandre Bilodeau
DOI:10.1016/j.imavis.2022.104612delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Video prediction is a challenging computer vision task that has a wide range of applications. In this work, we present a new family of Transformer-based models for video prediction. Firstly, an efficient local spatial- temporal separation attention mechanism is proposed to reduce the complexity of standard Transformers. Then, a full autoregressive model, a partial autoregressive model and a non-autoregressive model are developed based on the new efficient Transformer. The partial autoregressive model has a similar performance with the full autoregressive model but a faster inference speed. The non-autoregressive model not only achieves a faster infer-ence speed but also mitigates the quality degradation problem of the autoregressive counterparts, but it requires additional parameters and loss function for learning. Given the same attention mechanism, we conducted a com-prehensive study to compare the proposed three video prediction variants. Experiments show that the proposed video prediction models are competitive with more complex state-of-the-art convolutional-LSTM based models. The source code is available at https://github.com/XiYe20/VPTR.(c) 2022 Elsevier B.V. All rights reserved.
Keywords:
Video prediction
Transformers
Video representation learning
Autoregressive generative models
Non-autoregressive generative models
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

U
universite de montreal
Scholars:
4.6W
Papers: 3.8W
Citations: 46