Return
Split Computing for Mobile Devices: Energy and Latency Perspective
DOI:10.1109/TSC.2025.3564885.png)
Abstract
En 中文
To tackle the difficulties of running sophisticated deep neural network (DNN) models on mobile devices, split computing presents a viable solution by offloading computations to the edge server. Current split computing schemes typically aim to lower either inference latency or energy use separately; however, optimizing both simultaneously is quite challenging due to numerous shifting factors, such as intensive continuous DNN model inferences, DNN model traits, and device/network conditions. Moreover, in practical applications, edge server overload might lead to substantial queuing delays, adding complexity to the optimization process. This article outlines a joint optimization problem that simultaneously seeks to minimize both inference latency and energy consumption, with a distinct inclusion of queue clearance latency for an accurate analysis of the continuously generated DNN model inferences. To address this intricate optimization challenge, we introduce a low-complexity heuristic algorithm that sets split point decisions based on the residual energy of mobile devices for each DNN inference cycle. Upon evaluation, our proposed algorithm demonstrates notable improvements by reducing inference latency by between 73.37% and 99.39%, and cutting down energy usage by between 39.97% and 94.67% compared to fully local processing on mobile devices.
Keywords:
Split computing
deep neural network
mobile edge computing
mixed-integer nonlinear programming (MINLP)
joint optimization
sorting policy
Journal
IF:
5.8
Papers:
2.1K
Citations:
6.5K

