Return
Edge Delayed Deep Deterministic Policy Gradient: Efficient Continuous Control for Edge Scenarios
DOI:10.1109/TASE.2025.3604290.png)
Abstract
En 中文
Deep Reinforcement Learning (DRL) has emerged as a powerful paradigm for learning complex policies directly from high-dimensional input spaces, enabling advances across a variety of domains. Modern DRL algorithms often rely on dual-network Q-learning architectures to approximate optimal policies to overcome overestimation bias. Recent research has introduced approaches leveraging multiple Q-functions to further mitigate overestimation effects and enhance policy reliability. However, there is a growing emphasis on deploying DRL in edge scenarios, where privacy concerns and stringent hardware constraints necessitate highly efficient algorithms. In such environments, the computational and memory efficiency of learning methods is of critical importance. In this context, we propose Edge Delayed Deep Deterministic Policy Gradient (EdgeD3), a novel reinforcement learning algorithm specifically designed for edge computing settings. EdgeD3 offers significant reductions in GPU time (by 25%) and computational and memory usage (by 30%), while consistently achieving or surpassing the performance of state-of-the-art algorithms across multiple benchmarks and in real-world tasks. Note to Practitioners— Driven by the growing need for efficient computational solutions in automation, this research introduces the Edge Delayed Deep Deterministic Policy Gradient (EdgeD3), a novel Deep Reinforcement Learning algorithm designed for applications on edge devices with a limited computational budget. EdgeD3 is developed to enable on-device execution of policy learning, when computing resources are at a premium, such as in smart manufacturing and autonomous vehicle systems. EdgeD3 can learn highly effective policies, on par of state-of-the-art algorithms, while utilizing significantly fewer resources. This would empower autonomous devices to operate independently of cloud-based systems, fostering faster operational speeds and enhancing data privacy.
Keywords:
Continuous control
deep deterministic policy gradient
deep reinforcement learning
edge computing
Q-learning
Journal
IF:
6.4
Papers:
4.9K
Citations:
1.6W

