Return
PASS: A Priority-Based Model Assignment for Intelligent Application Acceleration in Edge Cloud
DOI:10.1109/JIOT.2025.3602068.png)
Abstract
En 中文
Thanks to the fine-grained resource management capabilities, serverless computing has been extended to edge cloud environments to support diverse Artificial Intelligence of Things (AIoT) applications, particularly those involving complex workflows of interdependent deep neural network (DNN) inference tasks. However, the inherent on-demand provisioning nature of serverless computing imposes the fact that, in serverless inference processes, the DNN models are typically maintained in the remote storage cluster and retrieved as needed. This inevitably incurs substantial latency overhead, particularly in resource-constrained edge cloud. In this article, we investigate how to accelerate the artificial intelligence (AI) application with joint consideration of both the model downloading time and intermediate data transmission time. We first formulate this problem into a nonlinear optimization form and prove it as NP-hard. We further propose a priority-based model assignment (PASS) algorithm and theoretically analyze its upper bound. The trace-driven experimental results demonstrate that our proposed algorithm outperforms other sate-of-art solutions and reduces the average application completion time by 23.6%.
Keywords:
Edge intelligence
Internet of Things (IoT)
model assignment
serverless computing
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

