Return
Sequence Modeling Architectures: Foundations [Special Issue on the Mathematics of Deep Learning]
F
A
R
A
C
J
F
A
V
DOI:10.1109/MSP.2026.3682808.png)
Abstract
En 中文
With the success of deep learning in the past decade, neural networks transitioned from simple layer stacking to sophisticated, highly complex designs. This article provides an overview of such deep learning architectures for sequence modeling, with a focus on their core mathematical principles: 1) we review the evolution of common sequence models from recurrent neural networks (RNNs), to Transformers, and modern state-space models (SSMs), exploring their foundational components such as gating, attention, and linear dynamical systems mechanisms. 2) We present general linear sequence mixers as a unifying framework to highlight the relationship between architectures and compare them in terms of inductive biases and algorithmic complexity. 3) We examine how practical considerations such as parallelizability and hardware awareness dictate their success. Finally, we discuss their impact on diverse fields of signal processing, including time-series analysis, natural language, and computer vision.
Keywords:
Deep learning
Sequential analysis
Artificial neural networks
Recurrent neural networks
Transformers
State-space methods
Training
Large scale integration
Signal processing algorithms
Journal
IF:
9.6
Papers:
1.1W
Citations:
1.7W
