Return
Deep deterministic policy gradient-model-agnostic meta-learning framework: Efficient adaptation in continuous control tasks
DOI:10.1016/j.rineng.2025.105139.png)
Abstract
En 中文
Deep reinforcement learning (DRL) demonstrates superior performance in continuous control tasks. However, extensive training across a variety of environments is frequently necessitates extensive training. This manuscript presents Meta-DDPG-MAML, which combines the Deep Deterministic Policy Gradient (DDPG) methodology with Model-Agnostic Meta-Learning (MAML) to augment both adaptability and efficiency. By incorporating meta-learning principles into the DDPG actor-critic architecture, the proposed approach facilitates swift adaptation utilizing minimal data. Empirical results from benchmark environments-LunarLanderContinuous-v2, BipedalWalker-v3, and Pendulum-v1-reveal that Meta-DDPG-MAML consistently surpasses DDPG in several performance metrics. In the LunarLander, it attains maximum returns approaching 200, effectively doubling the performance of DDPG while concurrently stabilizing episode durations. In the BipedalWalker, it exceeds 250 returns by the 5000th episode, whereas the DDPG plateaus below 300. In the Pendulum, both methodologies stabilize around-200. However, the Meta-DDPG-MAML framework exhibits a more consistent critic loss, averaging-52.5 compared to the fluctuations observed with DDPG. Across all tasks, Meta-DDPG-MAML achieves superior episode return stability, maintains higher critic loss consistency, and improves policy robustness under varying conditions. Furthermore, Meta-DDPG-MAML accelerates convergence rates by 30%-50% in intricate environments while achieving returns that are 20%-40% greater, thereby underscoring its efficiency and adaptability. This framework highlights the potential of integrating MAML into DRL methodologies for practical applications requiring rapid learning and robust performance. This work significantly enhances the practical applicability of DRL by enabling rapid adaptation and robust performance in real-world continuous control tasks.
Keywords:
Deep deterministic policy gradient
Deep reinforcement learning
Model-agnostic meta-learning
Continuous control tasks
Hyperparameters
Optimization

