1
Return

A Survey of Dynamic Token Computation in Transformers: Taxonomy, Stability, and Budget-Aware Evaluation

delete2026-08-04
delete0
delete
OA
AI
A
Abuzar Khan
A
Ahmad Junaid
A
Abid Iqbal
G
Ghassan Husnain
W
Wonseop Shin
S
Sangsoon Lim
DOI:10.1109/access.2026.3720106delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Transformers have delivered strong performance across vision, language and multimodal tasks but their computational cost grows rapidly with token count, creating a major obstacle for real-time inference, edge deployment and other resource-constrained settings. A central challenge is that many input tokens contribute unequally to final predictions, yet conventional Transformer pipelines still process them uniformly, resulting in redundant computation, latency overhead and inefficient resource use. To address this problem, this survey reviews dynamic token computation in Transformers, an important research area that seeks to reduce unnecessary computation while preserving predictive quality. The survey includes token pruning and dropping, token merging and aggregation, conditional token routing and adaptive depth or early-exit strategies across Transformer-family architectures. Its objective is to provide a taxonomy-driven evaluation framework that unifies existing methods under common token-decision axes and assesses them through routing stability, budget reliability, overhead-aware latency and matched benchmark criteria. The synthesis shows that prior work can be grouped by operation type, decision design, training strategy and budget control. Although reported efficiency gains are often promising, they depend strongly on selection overheads, hardware settings and whether theoretical compute reduction translates into actual latency improvement. Recurring weaknesses include inconsistent evaluation protocols, limited reproducibility, unstable routing decisions and weak guarantees under perturbation or distribution shift. This survey contributes a structured vocabulary, a budget- and stability-aware evaluation framework, and a matched benchmarking protocol for analyzing whether dynamic token methods provide reliable and deployment-relevant efficiency gains.
Keywords:
Dynamic token computation
transformers
conditional computation
token pruning
efficient inference

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.7W
Citations:
29.4W

Organization

K
king faisal university
Scholars:
2.0K
Papers: 1.2K
Citations: 0
C
chung-ang university
Scholars:
1.0K
Papers: 469
Citations: 0
C
Cecos University of IT and Emerging Sciences
Scholars:
18
Papers: 7
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers