arrow
Return

Fast and fair split computing for accelerating deep neural network (DNN) inference

delete2025-02-01
delete0
delete
OA
AI
C
Cha, Dongju
J
Jaewook Lee
S
Sangheon Pack *
DOI:10.1016/j.icte.2024.09.013delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Conventional split computing approaches for AI models that generate large outputs suffer from long transmission and inference times. Due to the limited resources of the edge server and selfish MDs, some MDs cannot offload their tasks and sacrifice their performance. To address these issues, we formulate an optimization problem to determine one or two split points that minimize inference latency while ensuring fair offloading among MDs. Additionally, we devise a low-complexity heuristic algorithm called fast and fair split computing (F2SC). Evaluation results demonstrate that F2SC reduces inference time by 3.8% similar to 20.1% compared to the conventional approaches while maintaining fairness. (c) 2024 The Author(s). Published by Elsevier B.V. on behalf of The Korean Institute of Communications and Information Sciences. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keywords:
Split computing
Split point decision
Deep neural network inference
Jain's fairness index
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ICT Express cover
ICT Express
IF:
4.2
Papers:
988
Citations:
2.5K

Organization

L
LG Electronics
Scholars:
765
Papers: 629
Citations: 0