arrow
Return

Deep deterministic policy gradient-model-agnostic meta-learning framework: Efficient adaptation in continuous control tasks

delete2025-06-01
delete0
delete
OA
AI
E
Ebrahim Hamid Sumiea *
S
Said Jadid Abdulkadir
H
Hitham Alhussian
S
Safwan Mahmood Al-Selwi
A
Alawi Alqushaibi
M
Mohammed Gamal Ragab
A
Abdul Muiz Fayyaz
DOI:10.1016/j.rineng.2025.105139delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep reinforcement learning (DRL) demonstrates superior performance in continuous control tasks. However, extensive training across a variety of environments is frequently necessitates extensive training. This manuscript presents Meta-DDPG-MAML, which combines the Deep Deterministic Policy Gradient (DDPG) methodology with Model-Agnostic Meta-Learning (MAML) to augment both adaptability and efficiency. By incorporating meta-learning principles into the DDPG actor-critic architecture, the proposed approach facilitates swift adaptation utilizing minimal data. Empirical results from benchmark environments-LunarLanderContinuous-v2, BipedalWalker-v3, and Pendulum-v1-reveal that Meta-DDPG-MAML consistently surpasses DDPG in several performance metrics. In the LunarLander, it attains maximum returns approaching 200, effectively doubling the performance of DDPG while concurrently stabilizing episode durations. In the BipedalWalker, it exceeds 250 returns by the 5000th episode, whereas the DDPG plateaus below 300. In the Pendulum, both methodologies stabilize around-200. However, the Meta-DDPG-MAML framework exhibits a more consistent critic loss, averaging-52.5 compared to the fluctuations observed with DDPG. Across all tasks, Meta-DDPG-MAML achieves superior episode return stability, maintains higher critic loss consistency, and improves policy robustness under varying conditions. Furthermore, Meta-DDPG-MAML accelerates convergence rates by 30%-50% in intricate environments while achieving returns that are 20%-40% greater, thereby underscoring its efficiency and adaptability. This framework highlights the potential of integrating MAML into DRL methodologies for practical applications requiring rapid learning and robust performance. This work significantly enhances the practical applicability of DRL by enabling rapid adaptation and robust performance in real-world continuous control tasks.
Keywords:
Deep deterministic policy gradient
Deep reinforcement learning
Model-agnostic meta-learning
Continuous control tasks
Hyperparameters
Optimization

Journal

Results in Engineering cover
Results in Engineering
IF:
7.9
Papers:
1.1W
Citations:
1.7W

Organization

U
Univ Teknol PETRONAS
Scholars:
235
Papers: 118
Citations: 45