返回
Multi-Model Running Latency Optimization in an Edge Computing Paradigm
DOI:10.3390/s22166097.png)
摘要
En 中文
Recent advances in both lightweight deep learning algorithms and edge computing increasingly enable multiple model inference tasks to be conducted concurrently on resource-constrained edge devices, allowing us to achieve one goal collaboratively rather than getting high quality in each standalone task. However, the high overall running latency for performing multi-model inferences always negatively affects the real-time applications. To combat latency, the algorithms should be optimized to minimize the latency for multi-model deployment without compromising the safety-critical situation. This work focuses on the real-time task scheduling strategy for multi-model deployment and investigating the model inference using an open neural network exchange (ONNX) runtime engine. Then, an application deployment strategy is proposed based on the container technology and inference tasks are scheduled to different containers based on the scheduling strategies. Experimental results show that the proposed solution is able to significantly reduce the overall running latency in real-time applications.
Keyword:
edge computing
latency optimization
multi-model
task scheduling
autonomous driving
AI
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.5
论文数:
7.2W
被引数:
20.9W
机构
引用论文
A comparative effectiveness trial of two faecal immunochemical tests for haemoglobin (FIT). Assessment of test performance and adherence in a single round of a population-based screening programme for colorectal cancer
Gut
IF0
Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing边缘智能: 用边缘计算铺平人工智能的最后一英里
PROCEEDINGS OF THE IEEE
IF25.9
Edge and Fog Computing Platform for Data Fusion of Complex Heterogeneous Sensors面向复杂异构传感器数据融合的边缘与雾计算平台
SENSORS
IF3.5
SIRT1 Overexpression Antagonizes Cellular Senescence with Activated ERK/S6k1 Signaling in Human Diploid Fibroblasts
PLoS ONE
IF0

