arrow
Return

Speculative Decoding on the SN40L Reconfigurable Dataflow Unit

delete2025-09-01
delete1
PRE
AI
D
Darshan Gandhi *
P
Pushkar Nandkar
N
Nasim Farahini
H
Håkan Zeffer
J
John J. Long
S
Samuel Rydh
M
Matheen Musaddiq
T
Tuowen Zhao
J
Joshua Brot
R
Reid Goodbar
Y
Yun Du
M
Mingran Wang
R
Raghu Prabhakar
DOI:10.1109/MM.2025.3592570delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Speculative decoding has emerged as a promising optimization strategy to accelerate generative artificial intelligence (AI) inference. This technique enhances the autoregressive decoding phase by using a smaller draft model to generate a few tokens, which are then validated by a larger model. However, synchronization overheads on both GPUs and hosts limit the performance gains. This article explores the implementation and optimization potential of batched speculative decoding within SambaNova's SN40L reconfigurable dataflow unit. We achieve more than 75% of the theoretical maximum performance for speculative decoding and delivering a 6x speedup compared to the baseline model. Furthermore, we demonstrate that a single SN40L rack, containing 16 sockets and offering high bandwidth memory bandwidth comparable to the DGX H100, outperforms the latter by up to 1.7x. The techniques and models discussed are deployed in SambaNova's production AI inference cloud, cloud.sambanova.ai, showcasing their tangible impact on large-scale AI applications.
Keywords:
Decoding
Computational modeling
Artificial intelligence
Training
Synchronization
Computer architecture
Throughput
Sockets
Runtime
Production

Journal

IEEE Micro cover
IEEE Micro
IF:
2.9
Papers:
125
Citations:
2.7K

Organization

No organization information available