arrow
返回

Cooperative dual-actor proximal policy optimization algorithm for multi-robot complex control task

delete2025-01-01
delete0
PRE
AI
J
Jacky Baltes
I
Ilham Akbar
S
Saeed Saeedvand *
DOI:10.1016/j.aei.2024.102960delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper introduces a novel multi-agent Deep Reinforcement Learning (DRL) framework named the Cooperative Dual-Actor Proximal Policy Optimization (CDA-PPO) algorithm, designed to address complex humanoid robot cooperative learning control tasks. Effective cooperation among multiple humanoid robots, particularly in scenarios involving complex walking gait control and external disturbances in dynamic environments, is a critical challenge. This is especially pertinent for tasks requiring precise coordination and control, such as joint object transportation. In various real-life scenarios, humanoid robots might need to cooperate to carry large objects in many scenarios. This capability is crucial for logistics, manufacturing, intelligent transportation, and search-and-rescue missions applications. Humanoid robots have gained significant popularity, and their use in these cooperative tasks is becoming more common. To address this challenge, we propose CDA-PPO, which introduces a learning-based communication platform between agents and employs two distinct policy networks for each agent. This dual-policy approach enhances the robots' ability to adapt to complex interactions and maintain stability while performing intricate tasks. We demonstrate the efficacy of CDA-PPO in a cooperative object-transportation scenario, where two humanoid robots collaborate to carry a table. The experimental results show that CDA-PPO significantly outperforms traditional methods, such as Independent PPO (IPPO), Multi-Agent PPO (MAPPO), and Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3), in terms of training efficiency, stability, reward acquisition, and humanoid robot cooperative balance control with effective coordination between robots. The findings underscore the potential of CDA-PPO to advance the field of cooperative multi-agent control problems, proposing the way for future research in complex robotics applications.
Keyword:
Deep reinforcement learning
Proximal policy optimization
Humanoid robotics
Isaac gym

期刊

Advanced Engineering Informatics 封面图
Advanced Engineering Informatics
IF:
9.9
论文数:
4.1K
被引数:
1.7W

机构

N
National Taiwan Normal University
学者数:
4.8K
论文数: 4.7K
被引数: 4.4K
引用论文

引用论文

err分享
err收藏
Circularly Polarized Printed Helix Antenna for L- and S-Bands Applications
err2020-04-14
err0
errOAAI
errA. Siahcheshm; J. Nourinia; Ch. Ghobadi; M. Shokri
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
miR-132, an experience-dependent microRNA, is essential for visual cortex plasticity
err2011-09-04
err0
errOAAI
errNikolaos Mellios; Hiroki Sugihara; Jorge Castro; Abhishek Banerjee; Chuong Le; Arooshi Kumar; Benjamin Crawford; Julia Strathmann; Daniela Tropea; Stuart S Levine; Dieter Edbauer; Mriganka Sur
err分享
err收藏
学者 查看更多内容