arrow
返回

Batch Prioritization in Multigoal Reinforcement Learning

delete2020-01-01
delete8
delete
OA
AI
L
Luiz Felipe Vecchietti
K
Kim, Taeyoung
K
Kyujin Choi
J
Junhee Hong
D
Dongsoo Har *
DOI:10.1109/ACCESS.2020.3012204delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
In multigoal reinforcement learning, an agent interacts with an environment and learns to achieve multiple goals. The goal-conditioned policy is trained to effectively generalize its behavior for multiple goals. During training, the experiences collected by the agent are randomly sampled from a replay buffer. Because biased sampling of achieved goals affects the success rate of a given task, it should be avoided by considering the valid goal space, introduced here as the set of goals to achieve, and the current competence of the policy. To this end, a novel prioritization method for creation of batches, e.g., collections of samples, is proposed. Candidate batches are sampled and associated with costs; in each iteration the batch with the minimum cost is chosen to train the policy. The cost function is modeled by an intended goal, which is proposed as a hypothetical goal that the policy is trying to learn in each cycle, and the information of the valid goal space. The minimum cost of the batch selected for each iteration decreases throughout training as the policy learns to achieve goals near the center of the valid goal space. The proposed batch prioritization method is combined with hindsight experience replay (HER) for experiments in robotic control tasks presented in the OpenAI gym suite to demonstrate learning performance comparable to that of other state-of-the-art prioritization methods. As a result, the proposed batch prioritization method can achieve improved learning performance in 4 out of 5 tasks, particularly for harder tasks. The experimental results suggest that the proposed method for the creation of training batches, using the valid goal space information and current competence of the policy, can enhance learning performance in multigoal tasks with high-dimensional goal space.
Keyword:
Training
Task analysis
Learning (artificial intelligence)
Robots
Erbium
Cost function
Aerospace electronics
Experience replay
batch prioritization
goal distribution
reinforcement learning
intended goal
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

G
Gachon University
学者数:
8.2K
论文数: 9.3K
被引数: 8.6K
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Uncoupling protein-2 polymorphisms in type 2 diabetes, obesity, and insulin secretion
err2004-01-01
err0
PREAI
errHua Wang; Winston S. Chu; Tong Lu; Sandra J. Hasstedt; Philip A. Kern; Steven C. Elbein
err分享
err收藏
Sampling Rate Decay in Hindsight Experience Replay for Robot Control
err2022-03-01
err25
PREAI
errVecchietti, Luiz Felipe; Seo, Minah; Har, Dongsoo
err分享
err收藏
学者 查看更多内容