arrow
Return

ESTSformer: Efficient spatio-temporal spiking transformer

delete2025-07-02
delete0
PRE
AI
C
Chengzhuo Lu
H
Huilin Du
W
Wenjie Wei
Q
Qian Sun
Y
Yuchen Wang
D
Dingyi Zeng
W
Wenyu Chen
M
Malu Zhang
Y
Yang Yang
DOI:10.1016/j.neunet.2025.107786delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Bio-inspired Spiking Neural Networks (SNNs) have garnered significant attention for their binary, asynchronous, and event-driven computing. These characteristics make SNNs a compelling alternative to traditional Artificial Neural Networks (ANNs). As Transformers expand across AI domains, integrating them with SNNs offers a promising path to high-performance and efficient models. A critical challenge in this area is effectively leveraging SNNs’ inherent spatio-temporal advantages. Unfortunately, existing studies on spiking spatio-temporal attention mechanisms exhibit certain limitations. In this paper, we first analyze the drawbacks of vanilla spatiotemporal self-attention (STSA), specifically its substantial computational and storage demands that escalate with increasing time steps. To address this, we propose an efficient spatiotemporal self-attention (ESTSA) mechanism. ESTSA divides attention heads into distinct sets for temporal and spatial information extraction, a simple yet effective division that significantly reduces both computational and storage overhead. Based on the ESTSA, we construct an efficient spiking transformer architecture, termed ESTSformer, which modifies residual connections within its encoder modules to ensure purely spike-driven computation throughout the network. Extensive experiments on both neuromorphic and static datasets demonstrate that our method achieves superior performance and efficiency compared to advanced existing works.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

No organization information available