arrow
Return

Edge Delayed Deep Deterministic Policy Gradient: Efficient Continuous Control for Edge Scenarios

delete2025-01-01
delete0
PRE
AI
A
Alberto Sinigaglia
N
Niccolò Turcato
R
Ruggero Carli
G
Gian Antonio Susto
DOI:10.1109/TASE.2025.3604290delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep Reinforcement Learning (DRL) has emerged as a powerful paradigm for learning complex policies directly from high-dimensional input spaces, enabling advances across a variety of domains. Modern DRL algorithms often rely on dual-network Q-learning architectures to approximate optimal policies to overcome overestimation bias. Recent research has introduced approaches leveraging multiple Q-functions to further mitigate overestimation effects and enhance policy reliability. However, there is a growing emphasis on deploying DRL in edge scenarios, where privacy concerns and stringent hardware constraints necessitate highly efficient algorithms. In such environments, the computational and memory efficiency of learning methods is of critical importance. In this context, we propose Edge Delayed Deep Deterministic Policy Gradient (EdgeD3), a novel reinforcement learning algorithm specifically designed for edge computing settings. EdgeD3 offers significant reductions in GPU time (by 25%) and computational and memory usage (by 30%), while consistently achieving or surpassing the performance of state-of-the-art algorithms across multiple benchmarks and in real-world tasks. Note to Practitioners— Driven by the growing need for efficient computational solutions in automation, this research introduces the Edge Delayed Deep Deterministic Policy Gradient (EdgeD3), a novel Deep Reinforcement Learning algorithm designed for applications on edge devices with a limited computational budget. EdgeD3 is developed to enable on-device execution of policy learning, when computing resources are at a premium, such as in smart manufacturing and autonomous vehicle systems. EdgeD3 can learn highly effective policies, on par of state-of-the-art algorithms, while utilizing significantly fewer resources. This would empower autonomous devices to operate independently of cloud-based systems, fostering faster operational speeds and enhancing data privacy.
Keywords:
Continuous control
deep deterministic policy gradient
deep reinforcement learning
edge computing
Q-learning

Journal

IEEE Transactions on Automation Science and Engineering cover
IEEE Transactions on Automation Science and Engineering
IF:
6.4
Papers:
4.9K
Citations:
1.6W

Organization

U
university of padova
Scholars:
3.1K
Papers: 1.4K
Citations: 1