arrow
Return

Large Language Models (LLMs) Inference Offloading and Resource Allocation in Cloud-Edge Computing: An Active Inference Approach

delete2024-12-01
delete1
PRE
AI
Y
Ying He
J
Jingcheng Fang
F
F. Richard Yu *
V
Victor C. M. Leung
DOI:10.1109/TMC.2024.3415661delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the increasing popularity and demands for large language model applications on mobile devices, it is difficult for resource-limited mobile terminals to run large-model inference tasks efficiently. Traditional deep reinforcement learning (DRL) based approaches have been used to offload large language models (LLMs) inference tasks to servers. However, existing DRL solutions suffer from data inefficiency, insensitivity to latency requirements, and non-adaptability to task load variations, which will degrade the performance of LLMs. In this paper, we propose a novel approach based on active inference for LLMs inference task offloading and resource allocation in cloud-edge computing. Extensive simulation results show that our proposed method has superior performance over mainstream DRLs, improves in data utilization efficiency, and is more adaptable to changing task load scenarios.
Keywords:
Task analysis
Computational modeling
Cloud computing
Resource management
Edge computing
Artificial neural networks
Predictive models
Active inference
cloud-edge computing
large language model
reinforcement learning
resource allocation
task offloading

Journal

IEEE Transactions on Mobile Computing cover
IEEE Transactions on Mobile Computing
IF:
9.2
Papers:
5.6K
Citations:
1.8W

Organization

S
shenzhen university
Scholars:
4.5W
Papers: 3.4W
Citations: 72
U
University of British Columbia
Scholars:
7.0W
Papers: 6.1W
Citations: 8.6W