arrow
返回

A Universal Empirical Dynamic Programming Algorithm for Continuous State MDPs

delete2020-01-01
delete12
delete
OA
AI
W
William B. Haskell
R
Rahul Jain *
H
Hiteshi Sharma
P
P. L. Yu
DOI:10.1109/TAC.2019.2907414delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We propose universal randomized function approximation-based empirical value learning (EVL) algorithms for Markov decision processes. The empirical nature comes from each iteration being done empirically from samples available from simulations of the next state. This makes the Bellman operator a random operator. A parametric and a nonparametric method for function approximation using a parametric function space and a reproducing kernel Hilbert space respectively are then combined with EVL. Both function spaces have the universal function approximation property. Basis functions are picked randomly. Convergence analysis is performed using a random operator framework with techniques from the theory of stochastic dominance. Finite time sample complexity bounds are derived for both universal approximate dynamic programming algorithms. Numerical experiments support the versatility and computational tractability of this approach.
Keyword:
Approximation algorithms
Heuristic algorithms
Probabilistic logic
Dynamic programming
Function approximation
Convergence
Complexity theory
Continuous state-space Markov decision processes (MDPs)
dynamic programming (DP)
reinforcement learning (RL)
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

U
university of southern california
学者数:
4.7W
论文数: 3.8W
被引数: 51
N
National University of Singapore
学者数:
7.5W
论文数: 6.5W
被引数: 11.4W
引用论文

引用论文

Inside the Black Box
err
IF0
err2011-12-01
err0
PREAI
errRishi K Narang
err分享
err收藏
Similarity Measures and Dimensionality Reduction Techniques for Time Series Data Mining
err2012-09-12
err0
errOAAI
errCarmelo Cassisi; Placido Montalto; Marco Aliotta; Andrea Cannata; Alfredo Pulvirenti
err分享
err收藏
学者 查看更多内容