arrow
Return

A Universal Empirical Dynamic Programming Algorithm for Continuous State MDPs

delete2020-01-01
delete12
delete
OA
AI
W
William B. Haskell
R
Rahul Jain *
H
Hiteshi Sharma
P
P. L. Yu
DOI:10.1109/TAC.2019.2907414delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We propose universal randomized function approximation-based empirical value learning (EVL) algorithms for Markov decision processes. The empirical nature comes from each iteration being done empirically from samples available from simulations of the next state. This makes the Bellman operator a random operator. A parametric and a nonparametric method for function approximation using a parametric function space and a reproducing kernel Hilbert space respectively are then combined with EVL. Both function spaces have the universal function approximation property. Basis functions are picked randomly. Convergence analysis is performed using a random operator framework with techniques from the theory of stochastic dominance. Finite time sample complexity bounds are derived for both universal approximate dynamic programming algorithms. Numerical experiments support the versatility and computational tractability of this approach.
Keywords:
Approximation algorithms
Heuristic algorithms
Probabilistic logic
Dynamic programming
Function approximation
Convergence
Complexity theory
Continuous state-space Markov decision processes (MDPs)
dynamic programming (DP)
reinforcement learning (RL)
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Automatic Control cover
IEEE Transactions on Automatic Control
IF:
7
Papers:
1.3W
Citations:
6.7W

Organization

U
university of southern california
Scholars:
4.7W
Papers: 3.8W
Citations: 51
N
National University of Singapore
Scholars:
7.5W
Papers: 6.5W
Citations: 11.4W
Cited Papers

Cited Papers

DPABI: Data Processing & Analysis for (Resting-State) Brain Imaging
err2016-04-13
err0
PREAI
errChao-Gan Yan; Xin-Di Wang; Xi-Nian Zuo; Yu-Feng Zang
errShare
errSave
Inside the Black Box
err
IF0
err2011-12-01
err0
PREAI
errRishi K Narang
errShare
errSave
Similarity Measures and Dimensionality Reduction Techniques for Time Series Data Mining
err2012-09-12
err0
errOAAI
errCarmelo Cassisi; Placido Montalto; Marco Aliotta; Andrea Cannata; Alfredo Pulvirenti
errShare
errSave
researcher View more