arrow
Return

DVFO: Learning-Based DVFS for Energy-Efficient Edge-Cloud Collaborative Inference

delete2024-10-01
delete2
delete
OA
AI
Z
Ziyang Zhang
Y
Yang Zhao
李焕 (Huan Li)
C
Changyao Lin
刘洁 (Jie Liu) *
DOI:10.1109/TMC.2024.3357218delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Due to limited resources on edge and different characteristics of deep neural network (DNN) models, it is a big challenge to optimize DNN inference performance in terms of energy consumption and end-to-end latency. In addition to dynamic voltage frequency scaling (DVFS) technique, edge-cloud architecture provides a collaborative approach for efficient DNN inference. However, current edge-cloud collaborative inference methods have not optimized various compute resources on edge devices. Thus, we propose DVFO, a novel DVFS-enabled edge-cloud collaborative inference framework, which co-optimizes DVFS and offloading parameters via deep reinforcement learning (DRL). Specifically, DVFO automatically co-optimizes 1) the CPU, GPU and memory frequencies of edge devices, and 2) the offloaded feature map. In addition, it leverages a thinking-while-moving concurrent mechanism to accelerate the DRL learning process, and a spatial-channel attention mechanism to identify the less important DNN feature map for efficient offloading. This approach improves inference performance for different DNN models under various edge-cloud network conditions. Extensive evaluations using two datasets and six widely-deployed DNN models on five heterogeneous edge devices show that DVFO significantly reduces the energy consumption by 33% on average, compared to state-of-the-art schemes. Moreover, DVFO achieves up to 28.6%similar to 59.1% end-to-end latency reduction, while maintaining accuracy within 1% loss on average.
Keywords:
Graphics processing units
Collaboration
Energy consumption
Computational modeling
Performance evaluation
Artificial neural networks
Servers
Edge computing
DVFS technology
collaborative inference
deep reinforcement learning

Journal

IEEE Transactions on Mobile Computing cover
IEEE Transactions on Mobile Computing
IF:
9.2
Papers:
5.6K
Citations:
1.8W

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66