arrow
Return

Optimizing Multi-DNN Inference on Mobile Devices Through Heterogeneous Processor Co-Execution

delete2025-12-23
delete0
PRE
AI
Y
Yunquan Gao
Z
Zhiguo Zhang
P
Praveen Kumar Donta
C
Chinmaya Kumar Dehury
X
Xiujun Wang
D
Dusit Niyato
Q
Qiyang Zhang
DOI:10.1109/TMC.2025.3647031delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep Neural Networks (DNNs) are increasingly adopted across various industries, driving the demand for deploying their capabilities on mobile devices. However, current mobile inference frameworks often rely on a single processor to execute each model inference, limiting hardware utilization and leading to suboptimal performance and energy efficiency. Expanding DNN accessibility on mobile platforms requires more adaptive and resource-efficient solutions to meet increasing computational demands without compromising device functionality. Nevertheless, performing parallel inference of multiple DNNs on heterogeneous processors remains a significant challenge. Existing studies have explored partitioning DNN operations into subgraphs to enable parallel execution across heterogeneous processors. However, these approaches typically generate excessive subgraphs based solely on hardware compatibility, increasing scheduling complexity and memory management overhead. To address these limitations, we propose the Advanced Multi-DNN Model Scheduling (ADMS) strategy that optimizes multi-DNN inference across heterogeneous processors on mobile devices. ADMS constructs an offline subgraph partitioning strategy that considers both hardware support for operations and scheduling granularity. It also employs a processor-state-aware scheduling algorithm to dynamically balance workloads based on real-time system conditions. This ensures efficient workload distribution and maximizes the utilization of available processors. Experimental results demonstrate that, compared to vanilla inference frameworks, ADMS achieves a 4.04× reduction in multi-DNN inference latency.
Keywords:
Deep Neural Networks
mobile devices
heterogeneous processors
co-execution
multi-DNN inference

Journal

IEEE Transactions on Mobile Computing cover
IEEE Transactions on Mobile Computing
IF:
9.2
Papers:
5.6K
Citations:
1.8W

Organization

C
China Telecom Corporation Ltd.
Scholars:
10
Papers: 10
Citations: 0
S
stockholm university
Scholars:
1.8K
Papers: 1.0K
Citations: 0
U
university of tartu
Scholars:
1.3K
Papers: 559
Citations: 0
A
anhui university of technology
Scholars:
9.4K
Papers: 5.5K
Citations: 9
N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
P
peking university
Scholars:
11.7W
Papers: 8.7W
Citations: 146
researcher View more organizations