Return
Recursive Mamba: Fractal recursive state-space networks with global memory banks for efficient sequence modeling
M
A
S
A
DOI:10.1016/j.neucom.2026.134750.png)
Abstract
En 中文
The dominant paradigm in deep sequence modeling has long followed a doctrine of parameter proliferation: deeper networks demand proportionally more parameters, with each layer maintaining independent computational machinery and isolated state representations. This architectural philosophy, while effective, conflicts with fundamental principles of cognitive economy observed in biological neural systems, where resource sharing and compressed memory consolidation enable remarkable efficiency. We challenge this paradigm by introducing the Fractal Multi-Scale State-Space (FMSS) architecture, a novel framework that achieves competitive accuracy with contemporary Mamba-2 baselines while requiring only 24.4% of the original parameters—a compression ratio exceeding four-to-one. Our approach rests on three synergistic innovations: (i) a single-core shared SSM paradigm wherein one state-space module serves all layers through recursive fractal composition; (ii) a GlobalStateBank mechanism that compresses the high-dimensional SSM state by a factor of 64 × while preserving discriminative capacity through momentum-updated slot-based memory; and (iii) a multi-exit fractal framework with cross-exit contrastive supervision and progressive self-distillation that extracts maximal representational utility from the compressed backbone. Across three standard text classification benchmarks (AG News, IMDB, SST-2), FMSS demonstrates a mean accuracy improvement of +0.27% over the full-capacity Mamba-2 baseline while reducing parameters by 73.6% on average. On AG News specifically, FMSS achieves 92.53% accuracy compared to 91.57% for Mamba-2, representing nearly a full percentage point improvement with four times fewer parameters. These results suggest that the path toward parameter-efficient sequence modeling lies not in ever-larger architectures, but in principled compression guided by philosophies of cognitive economy and shared representation.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
