Return
CDAMetaRL: Continuous Dynamics Adaptive Meta Reinforcement Learning
J
C
C
DOI:10.1109/tmech.2025.3648420.png)
Abstract
En 中文
Modern deep reinforcement learning methods are commonly trained in the physics simulators for robot control tasks instead of on real hardware due to inefficiency and safety issues. However, the simulators hardly emulate physical work environment and complex dynamics changes. The incurred domain discrepancy negatively affects the policy’s performance in the real world. In this article, we propose a continuous dynamics adaptive meta-reinforcement learning (CDAMetaRL) framework to realize dynamic adaptation for flexible and intelligent control. CDAMetaRL introduces, first, an adversarial dynamic alignment method that involves actor and critic encoders and a continuous domain discriminator for learning domain-invariant features, and second, a meta-optimization method to further improve adversarial domain adaptation in meta-reinforcement learning. Extensive experiments show CDAMetaRL outperforms several state-of-the-art domain adaptive reinforcement learning methods on five classic sim-to-sim and one typical sim-to-real robot control benchmarks.
Keywords:
Adversarial domain adaptation
reinforcement learning (RL)
sim-to-real transfer
Journal
I
IF:
7.3
Papers:
5.4K
Citations:
2.4W
