arrow
返回

Seformer: a long sequence time-series forecasting model based on binary position encoding and information transfer regularization

delete2022-11-28
delete5
PRE
AI
P
Pengyu Zeng
X
Xiaofeng Zhou *
S
Shuai Li
刘鹏杰 封面图
刘鹏杰 (Pengjie Liu)
DOI:10.1007/s10489-022-04263-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Long sequence time-series forecasting (LSTF) problems, such as weather forecasting, stock market forecasting, and power resource management, are widespread in the real world. The LSTF problem requires a model with high prediction accuracy. Recent studies have shown that the transformer model architecture is the most promising model structure for LSTF problems compared with other model architectures. The transformer model has the property of permutation equivalence, which leads to the importance of sequence position encoding, an essential process in model training. Currently, the continuous dynamics models constructed for position encoding using the neural differential equations (neural ODEs) method can model sequence position information well. However, we have found that there are some limitations when neural ODEs are applied to the LSTF problem, including the time cost problem, the baseline drift problem, and the information loss problem; thus, neural ODEs cannot be directly applied to the LSTF problem. To address this problem, we design a binary position encoding-based regularization model for long sequence time-series prediction, named Seformer, which has the following structure: 1) The binary position encoding mechanism, including intrablock and interblock position encoding. For intrablock position encoding, we design a simple ODE method by discretizing the continuum dynamics model, which reduces the time cost required to compute neural ODEs while maintaining their dynamics properties to the maximum extent. In interblock position encoding, a chunked recursive form is adopted to alleviate the baseline drift problem caused by eigenvalue explosion. 2) Information transfer regularization mechanism: By regularizing the model intermediate hidden variables as well as the encoder-decoder connection variables, we can reduce information loss during the model training process while ensuring the smoothness of the position information. Extensive experimental results obtained on six large-scale datasets show a consistent improvement in our approach over the baselines.
Keyword:
Long sequence time-series forecasting
Transformer
Position encoding
Regularization method
Conditional variational autoencoder

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Wireless sensor network for AI-based flood disaster detection
err2020-08-07
err42
PREAI
errAl Qundus, Jamal; Dabbour, Kosai; Gupta, Shivam; Meissonier, Regis; Paschke, Adrian
err分享
err收藏
First Human Results With the 256 Channel Intelligent Micro Implant Eye (IMIE 256)
err2021-10-27
err0
errOAAI
errHuizhuo Xu; Xingwu Zhong; Changlin Pang; Jing Zou; Wangling Chen; Xianggui Wang; Shanxiang Li; Yuntao Hu; Didier S. Sagan; Philip T. Weiss; Yangyi Yao; Jiayi Xiang; Margot S. Dayan; Mark S. Humayun; Yu-Chong Tai
err分享
err收藏
Time series modelling to forecast the confirmed and recovered cases of COVID-19
err2020-09-01
err146
errOAAI
errMaleki, Mohsen; Mahmoudi, Mohammad Reza; Wraith, Darren; Pho, Kim-Hung
err分享
err收藏
学者 查看更多内容