arrow
返回

Learning to search: Functional gradient techniques for imitation learning

delete2009-06-17
delete148
PRE
AI
D
David Silver
DOI:10.1007/s10514-009-9121-3delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Programming robot behavior remains a challenging task. While it is often easy to abstractly define or even demonstrate a desired behavior, designing a controller that embodies the same behavior is difficult, time consuming, and ultimately expensive. The machine learning paradigm offers the promise of enabling programming by demonstration for developing high-performance robotic systems. Unfortunately, many behavioral cloning (Bain and Sammut in Machine intelligence agents. London: Oxford University Press, 1995; Pomerleau in Advances in neural information processing systems 1, 1989; LeCun et al. in Advances in neural information processing systems 18, 2006) approaches that utilize classical tools of supervised learning (e.g. decision trees, neural networks, or support vector machines) do not fit the needs of modern robotic systems. These systems are often built atop sophisticated planning algorithms that efficiently reason far into the future; consequently, ignoring these planning algorithms in lieu of a supervised learning approach often leads to myopic and poor-quality robot performance. While planning algorithms have shown success in many real-world applications ranging from legged locomotion (Chestnutt et al. in Proceedings of the IEEE-RAS international conference on humanoid robots, 2003) to outdoor unstructured navigation (Kelly et al. in Proceedings of the international symposium on experimental robotics (ISER), 2004; Stentz et al. in AUVSI's unmanned systems, 2007), such algorithms rely on fully specified cost functions that map sensor readings and environment models to quantifiable costs. Such cost functions are usually manually designed and programmed. Recently, a set of techniques has been developed that explore learning these functions from expert human demonstration. These algorithms apply an inverse optimal control approach to find a cost function for which planned behavior mimics an expert's demonstration. The work we present extends the Maximum Margin Planning (MMP) (Ratliff et al. in Twenty second international conference on machine learning (ICML06), 2006a) framework to admit learning of more powerful, non-linear cost functions. These algorithms, known collectively as LEARCH (LEArning to seaRCH), are simpler to implement than most existing methods, more efficient than previous attempts at non-linearization (Ratliff et al. in NIPS, 2006b), more naturally satisfy common constraints on the cost function, and better represent our prior beliefs about the function's form. We derive and discuss the framework both mathematically and intuitively, and demonstrate practical real-world performance with three applied case-studies including legged locomotion, grasp planning, and autonomous outdoor unstructured navigation. The latter study includes hundreds of kilometers of autonomous traversal through complex natural environments. These case-studies address key challenges in applying the algorithm in practical settings that utilize state-of-the-art planners, and which may be constrained by efficiency requirements and imperfect expert demonstration.
Keyword:
Imitation learning
Structured prediction
Subgradient methods
Nonparametric optimization
Functional gradient techniques
Robotics
Planning
Autonomous navigation
Quadrupedal locomotion
Grasping

期刊

Autonomous Robots 封面图
Autonomous Robots
IF:
4.3
论文数:
1.7K
被引数:
5.0K

机构

C
Carnegie Mellon University
学者数:
1.4W
论文数: 1.4W
被引数: 2.7W
引用论文

引用论文

Disentangling in vivo the effects of iron content and atrophy on the ageing human brain
err2014-12-01
err0
errOAAI
errS. Lorio; A. Lutti; F. Kherif; A. Ruef; J. Dukart; R. Chowdhury; R.S. Frackowiak; J. Ashburner; G. Helms; N. Weiskopf; B. Draganski
err分享
err收藏
DOES CONTINGENT VALUATION MEASURE PREFERENCES?
err2015-03-08
err0
PREAI
errPETER A. DIAMOND; JERRY A. HAUSMAN; GREGORY K. LEONARD
err分享
err收藏
Neuroblastoma cells inhibit the immunostimulatory function of dendritic cells
err2003-06-01
err0
PREAI
errXiao Chen; Kara Doffek; Sonia L Sugg; Joel Shilyansky
err分享
err收藏
Involvement of fos in spontaneous and ultraviolet light‐induced genetic changes
err2006-07-19
err0
PREAI
errSusanne van den Berg; Bernd Kaina; Hans J. Rahmsdorf; Helmut Ponta; Peter Herrlich
err分享
err收藏
Chemical and Kinetic Reaction Mechanisms of Quinohemoprotein Amine Dehydrogenase fromParacoccus denitrificans
err2003-08-26
err0
PREAI
errDapeng Sun; Kazutoshi Ono; Toshihide Okajima; Katsuyuki Tanizawa; Mayumi Uchida; Yukio Yamamoto; F. Scott Mathews; Victor L. Davidson
err分享
err收藏
学者 查看更多内容