Return
MoR-ST: content-adaptive recursive computation for efficient vision transformers
DOI:10.1007/s00371-026-04713-6.png)
Abstract
En 中文
Vision transformer models have achieved remarkable success in computer vision, with Swin transformer serving as a versatile hierarchical backbone for classification, detection, and segmentation tasks. However, increasing model scales bring prohibitive memory consumption and computational costs that hinder deployment in resource-constrained environments. This paper proposes MoR-ST, a content-adaptive recursive computation framework that integrates the mixture of recursions (MoR) paradigm into the Swin transformer architecture. The core mechanism is an adaptive token-level recursion strategy: a lightweight routing network dynamically assigns recursion depths to individual tokens based on their semantic complexity, enabling simple tokens (e.g., background regions) to exit early while complex tokens (e.g., textured objects) undergo deeper processing. We evaluate MoR-ST on three standard benchmarks: ImageNet-1K/22K image classification, COCO object detection, and ADE20K semantic segmentation. Experimental results show that MoR-ST achieves competitive accuracy compared to the original Swin transformer while reducing parameters by approximately 50% and increasing inference speed by up to 2 $$\times $$ . On ImageNet-1K, MoR-Swin-T attains 80.5% Top-1 accuracy with only 15M parameters (vs. 28M for Swin-T). On COCO detection with Mask R-CNN, MoR-Swin-B achieves 51.9 box mAP with a 44M-parameter backbone (vs. 88M for Swin-B; the $$\sim $$ 50% reduction occurs in the backbone, with the same Mask R-CNN head architecture across models, up to backbone-dependent interface layers). Unlike token-pruning methods that irreversibly discard information along the spatial dimension, or mixture-of-experts models whose parameter count grows with the number of experts, MoR-ST operates along the depth dimension: It retains all tokens under a near-constant parameter budget and only varies the per-token recursion depth. This design naturally matches the spatial heterogeneity of real-world scenes—large simple-background regions alongside a few complex foreground objects—making it particularly well-suited to resource-constrained visual computing deployment. We emphasize that MoR-ST offers a favorable efficiency–accuracy trade-off rather than an accuracy gain: on ImageNet-1K, it stays within about 0.5 points of Swin-B while roughly halving the parameters and delivering up to 2 $$\times $$ speedup, and the value of this trade-off depends on the deployment regime, which we quantify rather than assert. These results demonstrate that adaptive recursive computation provides an effective pathway for building efficient vision backbones suitable for real-time visual computing applications.
Keywords:
Swin transformer
Mixture of recursions
Adaptive computation
Efficient vision backbone
Token-level recursion
Visual computing
Journal
IF:
2.9
Papers:
4.6K
Citations:
6.5K

