Return
ESTSformer: Efficient spatio-temporal spiking transformer
DOI:10.1016/j.neunet.2025.107786.png)
Abstract
En 中文
Bio-inspired Spiking Neural Networks (SNNs) have garnered significant attention for their binary, asynchronous, and event-driven computing. These characteristics make SNNs a compelling alternative to traditional Artificial Neural Networks (ANNs). As Transformers expand across AI domains, integrating them with SNNs offers a promising path to high-performance and efficient models. A critical challenge in this area is effectively leveraging SNNs’ inherent spatio-temporal advantages. Unfortunately, existing studies on spiking spatio-temporal attention mechanisms exhibit certain limitations. In this paper, we first analyze the drawbacks of vanilla spatiotemporal self-attention (STSA), specifically its substantial computational and storage demands that escalate with increasing time steps. To address this, we propose an efficient spatiotemporal self-attention (ESTSA) mechanism. ESTSA divides attention heads into distinct sets for temporal and spatial information extraction, a simple yet effective division that significantly reduces both computational and storage overhead. Based on the ESTSA, we construct an efficient spiking transformer architecture, termed ESTSformer, which modifies residual connections within its encoder modules to ensure purely spike-driven computation throughout the network. Extensive experiments on both neuromorphic and static datasets demonstrate that our method achieves superior performance and efficiency compared to advanced existing works.
Journal
IF:
6.3
Papers:
7.8K
Citations:
3.0W
Organization
No organization information available

