arrow
Return

MoR-ST: content-adaptive recursive computation for efficient vision transformers

delete2026-09-03
delete0
PRE
AI
Y
Yongbao Ai *
T
Tianxiang Gao
DOI:10.1007/s00371-026-04713-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Vision transformer models have achieved remarkable success in computer vision, with Swin transformer serving as a versatile hierarchical backbone for classification, detection, and segmentation tasks. However, increasing model scales bring prohibitive memory consumption and computational costs that hinder deployment in resource-constrained environments. This paper proposes MoR-ST, a content-adaptive recursive computation framework that integrates the mixture of recursions (MoR) paradigm into the Swin transformer architecture. The core mechanism is an adaptive token-level recursion strategy: a lightweight routing network dynamically assigns recursion depths to individual tokens based on their semantic complexity, enabling simple tokens (e.g., background regions) to exit early while complex tokens (e.g., textured objects) undergo deeper processing. We evaluate MoR-ST on three standard benchmarks: ImageNet-1K/22K image classification, COCO object detection, and ADE20K semantic segmentation. Experimental results show that MoR-ST achieves competitive accuracy compared to the original Swin transformer while reducing parameters by approximately 50% and increasing inference speed by up to 2 $$\times $$ . On ImageNet-1K, MoR-Swin-T attains 80.5% Top-1 accuracy with only 15M parameters (vs. 28M for Swin-T). On COCO detection with Mask R-CNN, MoR-Swin-B achieves 51.9 box mAP with a 44M-parameter backbone (vs. 88M for Swin-B; the $$\sim $$ 50% reduction occurs in the backbone, with the same Mask R-CNN head architecture across models, up to backbone-dependent interface layers). Unlike token-pruning methods that irreversibly discard information along the spatial dimension, or mixture-of-experts models whose parameter count grows with the number of experts, MoR-ST operates along the depth dimension: It retains all tokens under a near-constant parameter budget and only varies the per-token recursion depth. This design naturally matches the spatial heterogeneity of real-world scenes—large simple-background regions alongside a few complex foreground objects—making it particularly well-suited to resource-constrained visual computing deployment. We emphasize that MoR-ST offers a favorable efficiency–accuracy trade-off rather than an accuracy gain: on ImageNet-1K, it stays within about 0.5 points of Swin-B while roughly halving the parameters and delivering up to 2 $$\times $$ speedup, and the value of this trade-off depends on the deployment regime, which we quantify rather than assert. These results demonstrate that adaptive recursive computation provides an effective pathway for building efficient vision backbones suitable for real-time visual computing applications.
Keywords:
Swin transformer
Mixture of recursions
Adaptive computation
Efficient vision backbone
Token-level recursion
Visual computing

Journal

Visual Computer cover
Visual Computer
IF:
2.9
Papers:
4.6K
Citations:
6.5K

Organization

I
institute of military intelligence
Scholars:
3
Papers: 1
Citations: 0
Cited Papers

Cited Papers

Backpropagation Applied to Handwritten Zip Code Recognition
err1989-12-01
err0
PREAI
errY. LeCun; B. Boser; J. S. Denker; D. Henderson; R. E. Howard; W. Hubbard; L. D. Jackel
errShare
errSave
Lightweight Semantic Feature Extraction Model With Direction Awareness for Aerial Traffic Object Detection
err2026-04-01
err0
PREAI
errShen,Jiaquan; Liu,Ningzhong; Sun,Han; Wu,Shang; Liang,Zongzheng; Han,Lulu; Zhang,Yongxin; Li,Deguang
errShare
errSave
Adaptive Mixtures of Local Experts
err1991-02-01
err0
PREAI
errRobert A. Jacobs; Michael I. Jordan; Steven J. Nowlan; Geoffrey E. Hinton
errShare
errSave
An Anchor-Free Lightweight Deep Convolutional Network for Vehicle Detection in Aerial Images
err2022-12-01
err12
PREAI
errShen, Jiaquan; Zhou, Wangcheng; Liu, Ningzhong; Sun, Han; Li, Deguang; Zhang, Yongxin
errShare
errSave
PVT v2: Improved baselines with Pyramid Vision Transformer
err2022-09-01
err803
errOAAI
errWang, Wenhai; Xie, Enze; Li, Xiang; Fan, Deng-Ping; Song, Kaitao; Liang, Ding; Lu, Tong; Luo, Ping; Shao, Ling
errShare
errSave