返回
A masked autoencoder network for spatiotemporal predictive learning
DOI:10.1007/s10489-024-06214-2.png)
摘要
En 中文
This paper is about predictive learning, which is generating future frames given previous images. Suffering from the vanishing gradient problem, existing methods based on RNN and CNN can't capture the long-term dependencies effectively. To overcome the above dilemma, we present MastNet a spatiotemporal framework for long-term predictive learning. In this paper, we design a Transformer-based encoder-decoder with hierarchical structure. As for the transformer block, we adopt the spatiotemporal window based self-attention to reduce computational complexity and the spatiotemporal shifted window partitioning approach. More importantly, we build a spatiotemporal autoencoder by the random clip mask strategy, which leads to better feature mining for temporal dependencies and spatial correlations. Furthermore, we insert an auxiliary prediction head, which can help our model generate higher-quality frames. Experimental results show that the proposed MastNet achieves the best results in accuracy and long-term prediction on two spatiotemporal datasets compared with the state-of-the-art models.
Keyword:
Spatiotemporal prediction
Predictive learning
Hierarchical architecture
Spatiotemporal transformer blocks
Random clip masked autoencoder
Auxiliary head
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?
CAST: A convolutional attention spatiotemporal network for predictive learning
APPLIED INTELLIGENCE
IF3.5
Synthesis of n-type semiconducting diamond film using diphosphorus pentaoxide as the doping source以五氧化二磷为掺杂源合成n型半导体金刚石膜
Structural study of lanthanides(III) in aqueous nitrate and chloride solutions by EXAFS通过EXAFS对硝酸盐和氯化物水溶液中镧系元素 (III) 的结构研究

