arrow
Return

Game-Based LLM Inference Task Offloading for Edge Computing System

delete2026-03-23
delete0
PRE
AI
S
Shoulu Hou
甘敏 (Min Gan)
W
Wei Ni
Z
Zhongyi Zhai
X
Xiulei Liu
DOI:10.1109/TGCN.2026.3676823delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large Language Models (LLMs), with their powerful capabilities, are fundamentally transforming society. Cloud-based LLM deployment has drawbacks, including latency, lack of offline functionality, and high long-term costs. Edge computing-based LLM solutions address these issues by offloading inference tasks to edge servers. This paper formulates and optimizes an LLM inference framework that jointly accounts for inference time and predictive quality. Specifically, to address the natural tendency of selfish users for personal utility maximization, we develop a Game-theoretic Offloading Algorithm (GOALIT) to optimize LLM inference offloading. The approach enables distributed users to iteratively adjust their strategies and converge to a Nash equilibrium. Compared to the optimal tree-based search (OT-GAH), the proposed approach yields a 27% increase in token throughput and shortens inference latency by 20%, while maintaining lower perplexity under dynamic system loads. These findings confirm the effectiveness of our approach in resource-constrained edge environments.
Keywords:
Edge computing
LLM inference task offloading
quantization
game theory

Journal

I
IEEE Transactions on Green Communications and Networking
IF:
6.7
Papers:
1.3K
Citations:
4.3K

Organization

B
Beijing Information Science and Technology University
Scholars:
628
Papers: 315
Citations: 1.5K
E
edith cowan university
Scholars:
1.1K
Papers: 595
Citations: 1
G
guilin university of electronic technology
Scholars:
2.2K
Papers: 730
Citations: 0
researcher View more organizations