arrow
Return

FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding

delete2026-04-13
delete0
PRE
AI
Y
Yuchen Li
R
Rui Kong
Z
Zhonghao Lyu
Q
Qiyang Li
X
Xinran Chen
H
Hengyi Cai
L
Lingyong Yan
S
Shuaiqiang Wang
J
Jiashu Zhao
G
Guangxu Zhu
L
Linghe Kong
G
Guihai Chen
H
Haoyi Xiong
D
D. H.-L. Yin
DOI:10.1109/tmc.2026.3683574delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deploying large language models (LLMs) in mobile and edge computing environments is constrained by limited on-device resources, scarce wireless bandwidth, and frequent model evolution. Although edge-cloud collaborative inference with speculative decoding (SD) can reduce end-to-end latency by executing a lightweight draft model at the edge and verifying it with a cloud-side target model, existing frameworks fundamentally rely on tight coupling between the two models. Consequently, repeated model synchronization introduces excessive communication overhead, increasing end-to-end latency, and ultimately limiting the scalability of SD in edge environments. To address these limitations, we propose FlexSpec, a communication-efficient collaborative inference framework tailored for evolving edge-cloud systems. The core design of FlexSpec is a shared-backbone architecture that allows a single and static edge-side draft model to remain compatible with a large family of evolving cloud-side target models. By decoupling edge deployment from cloud-side model updates, FlexSpec eliminates the need for edge-side retraining or repeated model downloads, substantially reducing communication and maintenance costs. Furthermore, to accommodate time-varying wireless conditions and heterogeneous device constraints, we develop a channel-aware adaptive speculation mechanism that dynamically adjusts the speculative draft length based on real-time channel state information and device energy budgets. Extensive experiments demonstrate that FlexSpec achieves superior performance compared to conventional SD approaches in terms of inference efficiency.
Keywords:
Mobile computing
edge-cloud collaboration
large language models
speculative decoding
collaborative inference

Journal

IEEE Transactions on Mobile Computing cover
IEEE Transactions on Mobile Computing
IF:
9.2
Papers:
5.6K
Citations:
1.8W

Organization

B
baidu inc.
Scholars:
62
Papers: 26
Citations: 0
S
shanghai jiao tong university
Scholars:
15.6W
Papers: 11.6W
Citations: 159
S
Shenzhen Research Institute of Big Data
Scholars:
255
Papers: 351
Citations: 357
W
wilfrid university
Scholars:
2
Papers: 1
Citations: 0
K
kth royal institute of technology
Scholars:
864
Papers: 467
Citations: 0
researcher View more organizations