arrow
返回

Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm

delete2014-07-02
delete34
delete
OA
AI
R
Róbert Busa‐Fekete *
B
Balázs Szörényi
P
Paul Weng
程玮玮 (Weiwei Cheng)
E
Eyke Hüllermeier
DOI:10.1007/s10994-014-5458-8delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We introduce a novel approach to preference-based reinforcement learning, namely a preference-based variant of a direct policy search method based on evolutionary optimization. The core of our approach is a preference-based racing algorithm that selects the best among a given set of candidate policies with high probability. To this end, the algorithm operates on a suitable ordinal preference structure and only uses pairwise comparisons between sample rollouts of the policies. Embedding the racing algorithm in a rank-based evolutionary search procedure, we show that approximations of the so-called Smith set of optimal policies can be produced with certain theoretical guarantees. Apart from a formal performance and complexity analysis, we present first experimental studies showing that our approach performs well in practice.
Keyword:
Preference learning
Reinforcement learning
Evolutionary direct policy search
Racing algorithms
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

P
Philipps University Marburg
学者数:
1.3W
论文数: 1.0W
被引数: 10
S
Sorbonne Universite
学者数:
6.2W
论文数: 4.5W
被引数: 605