返回
Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks
DOI:10.1109/JSTSP.2025.3581478.png)
摘要
En 中文
Large language models (LLMs) have demonstrated remarkable success across various application domains, but their enormous sizes and computational demands pose significant challenges for deployment on resource-constrained edge devices. To address this issue, we propose a novel distributed on-device LLM inference framework that leverages tensor parallelism to partition the neural network tensors (e.g., weight matrices) of one LLM across multiple edge devices for collaborative inference. A key challenge in tensor parallelism is the frequent all-reduce operations for aggregating intermediate layer outputs across participating devices, which incurs significant communication overhead. To alleviate this bottleneck, we propose an over-the-air computation (AirComp) approach that harnesses the analog superposition property of wireless multiple-access channels to perform fast all-reduce steps. To utilize the heterogeneous computational capabilities of edge devices and mitigate communication distortions, we investigate a joint model assignment and transceiver optimization problem to minimize the average transmission error. The resulting mixed-timescale stochastic non-convex optimization problem is intractable, and we propose an efficient two-stage algorithm to solve it. Moreover, we prove that the proposed algorithm converges almost surely to a stationary point of the original problem. Comprehensive simulation results will show that the proposed framework outperforms existing benchmark schemes, achieving up to 5x inference speed acceleration and improving inference accuracy.
Keyword:
6G
distributed inference
large language models
over-the-air computation
tensor parallelism
期刊
IF:
13.7
论文数:
1.9K
被引数:
1.1W
机构
引用论文
Wirelessly Powered Data Aggregation for IoT via Over-the-Air Function Computation: Beamforming and Power Control通过空中功能计算实现物联网的无线供电数据聚合: 波束成形和功率控制
Embodied intelligence in manufacturing: leveraging large language models for autonomous industrial robotics制造业中的具体化智能: 利用大型语言模型实现自主工业机器人
Large Language Models (LLMs) Inference Offloading and Resource Allocation in Cloud-Edge Computing: An Active Inference Approach云边缘计算中的大型语言模型 (llm) 推理卸载和资源分配: 一种主动推理方法
Tackling Distribution Shifts in Task-Oriented Communication With Information Bottleneck用信息瓶颈解决任务导向通信中的分布偏移

