Return
A Survey of Dynamic Token Computation in Transformers: Taxonomy, Stability, and Budget-Aware Evaluation
A
A
A
G
W
S
DOI:10.1109/access.2026.3720106.png)
Abstract
En 中文
Transformers have delivered strong performance across vision, language and multimodal tasks but their computational cost grows rapidly with token count, creating a major obstacle for real-time inference, edge deployment and other resource-constrained settings. A central challenge is that many input tokens contribute unequally to final predictions, yet conventional Transformer pipelines still process them uniformly, resulting in redundant computation, latency overhead and inefficient resource use. To address this problem, this survey reviews dynamic token computation in Transformers, an important research area that seeks to reduce unnecessary computation while preserving predictive quality. The survey includes token pruning and dropping, token merging and aggregation, conditional token routing and adaptive depth or early-exit strategies across Transformer-family architectures. Its objective is to provide a taxonomy-driven evaluation framework that unifies existing methods under common token-decision axes and assesses them through routing stability, budget reliability, overhead-aware latency and matched benchmark criteria. The synthesis shows that prior work can be grouped by operation type, decision design, training strategy and budget control. Although reported efficiency gains are often promising, they depend strongly on selection overheads, hardware settings and whether theoretical compute reduction translates into actual latency improvement. Recurring weaknesses include inconsistent evaluation protocols, limited reproducibility, unstable routing decisions and weak guarantees under perturbation or distribution shift. This survey contributes a structured vocabulary, a budget- and stability-aware evaluation framework, and a matched benchmarking protocol for analyzing whether dynamic token methods provide reliable and deployment-relevant efficiency gains.
Keywords:
Dynamic token computation
transformers
conditional computation
token pruning
efficient inference
Journal
IF:
3.6
Papers:
9.7W
Citations:
29.4W
