arrow
返回

Supervised Meta-Reinforcement Learning With Trajectory Optimization for Manipulation Tasks

delete2024-04-01
delete3
delete
OA
AI
L
Lei Wang
张
张云洲 (Yunzhou Zhang) *
朱
朱德龙 (Delong Zhu)
S
Sonya Coleman
D
Dermot Kerr
DOI:10.1109/TCDS.2023.3286465delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Learning from small amounts of samples with reinforcement learning (RL) is challenging in many tasks, especially, in real-world applications, such as robotics. Meta-RL (meta-RL) has been proposed as an approach to address this problem by generalizing to new tasks through experience from previous similar tasks. However, these approaches generally perform meta-optimization by focusing direct policy search methods on validation samples from adapted policies, thus, requiring large amounts of on-policy samples during meta-training. To this end, we propose a novel algorithm called supervised meta-RL with trajectory optimization (SMRL-TO) by integrating model-agnostic meta-learning (MAML) and iterative LQR (iLQR)-based trajectory optimization. Our approach is designed to provide online supervision for validation samples through iLQR-based trajectory optimization and embed simple imitation learning into the meta-optimization rather than policy gradient steps. This is actually a bi-level optimization that needs to calculate several gradient updates in each meta-iteration, consisting of off-policy RL in the inner loop and online imitation learning in the outer loop. SMRL-TO can achieve significant improvements in sample efficiency without human-provided demonstrations, due to the effective supervision from iLQR-based trajectory optimization. In this article, we describe how to use iLQR-based trajectory optimization to obtain labeled data and then how leverage them to assist the training of meta-learner. Through a series of robotic manipulation tasks, we further show that compared with the previous methods, the proposed approach can substantially improve sample efficiency and achieve better asymptotic performance.
Keyword:
Task analysis
Trajectory optimization
Robots
Heuristic algorithms
Training
Complexity theory
Dynamical systems
Iterative LQR (iLQR)
meta learning
reinforcement learning (RL)
robotic manipulation
trajectory optimization

期刊

IEEE Transactions on Cognitive and Developmental Systems 封面图
IEEE Transactions on Cognitive and Developmental Systems
IF:
4.9
论文数:
1.0K
被引数:
3.5K

机构

U
Ulster University
学者数:
5.7K
论文数: 5.9K
被引数: 25
C
Chinese University of Hong Kong
学者数:
3.4W
论文数: 3.2W
被引数: 5.6W
N
northeastern university - china
学者数:
3.2W
论文数: 2.7W
被引数: 37
学者 查看更多机构
引用论文

引用论文

What Is the Relationship of Fear Avoidance to Physical Function and Pain Intensity in Injured Athletes?
err2018-02-16
err0
errOAAI
errStefan F. Fischerauer; Mojtaba Talaei-Khoei; Rens Bexkens; David C. Ring; Luke S. Oh; Ana-Maria Vranceanu
err分享
err收藏
Deep Reinforcement Learning for Autonomous Driving: A Survey用于自动驾驶的深度强化学习: 一项调查
err2022-06-01
err965
errOAAI
errKiran, B. Ravi; Sobh, Ibrahim; Talpaert, Victor; Mannion, Patrick; Al Sallab, Ahmad A.; Yogamani, Senthil; Perez, Patrick
err分享
err收藏
学者 查看更多内容