arrow
Return

Sequential Decision Making With Coherent Risk

delete2017-07-01
delete47
PRE
AI
A
Aviv Tamar *
Y
Yinlam Chow
M
Mohammad Ghavamzadeh
S
Shie Mannor
DOI:10.1109/TAC.2016.2644871delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We provide sampling-based algorithms for optimization under a coherent-risk objective. The class of coherent-risk measures is widely accepted in finance and operations research, among other fields, and encompasses popular risk-measures such as conditional value at risk and mean-semi-deviation. Our approach is suitable for problems in which tuneable parameters control the distribution of the cost, such as in reinforcement learning or approximate dynamic programming with a parameterized policy. Such problems cannot be solved using previous approaches. We consider both static risk measures and time-consistent dynamic risk measures. For static risk measures, our approach is in the spirit of policy gradient methods, while for the dynamic risk measures, we use actor-critic type algorithms.
Keywords:
Coherent risk
dynamic programming
Markov decision processes
policy gradient
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Automatic Control cover
IEEE Transactions on Automatic Control
IF:
7
Papers:
1.3W
Citations:
6.7W

Organization

U
University of California Berkeley
Scholars:
3.5W
Papers: 2.8W
Citations: 11.3W
S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W
I
Inria
Scholars:
3.5K
Papers: 2.5K
Citations: 343
University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K
researcher View more organizations