arrow
Return

Variable risk control via stochastic optimization

delete2013-07-01
delete30
PRE
AI
S
Scott Kuindersma *
R
Roderic A. Grupen
A
Andrew G. Barto
DOI:10.1177/0278364913476124delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We present new global and local policy search algorithms suitable for problems with policy-dependent cost variance (or risk), a property present in many robot control tasks. These algorithms exploit new techniques in non-parametric heteroscedastic regression to directly model the policy-dependent distribution of cost. For local search, the learned cost model can be used as a critic for performing risk-sensitive gradient descent. Alternatively, decision-theoretic criteria can be applied to globally select policies to balance exploration and exploitation in a principled way, or to perform greedy minimization with respect to various risk-sensitive criteria. This separation of learning and policy selection permits variable risk control, where risk-sensitivity can be flexibly adjusted and appropriate policies can be selected at runtime without relearning. We describe experiments in dynamic stabilization and manipulation with a mobile manipulator that demonstrate learning of flexible, risk-sensitive policies in very few trials.
Keywords:
policy search
Bayesian optimization
robot learning
risk-sensitive
dynamic mobile manipulation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

International Journal of Robotics Research cover
International Journal of Robotics Research
IF:
5
Papers:
2.4K
Citations:
1.5W

Organization

U
university of massachusetts system
Scholars:
3.8W
Papers: 3.5W
Citations: 42