返回
A Universal Empirical Dynamic Programming Algorithm for Continuous State MDPs
DOI:10.1109/TAC.2019.2907414.png)
摘要
En 中文
We propose universal randomized function approximation-based empirical value learning (EVL) algorithms for Markov decision processes. The empirical nature comes from each iteration being done empirically from samples available from simulations of the next state. This makes the Bellman operator a random operator. A parametric and a nonparametric method for function approximation using a parametric function space and a reproducing kernel Hilbert space respectively are then combined with EVL. Both function spaces have the universal function approximation property. Basis functions are picked randomly. Convergence analysis is performed using a random operator framework with techniques from the theory of stochastic dominance. Finite time sample complexity bounds are derived for both universal approximate dynamic programming algorithms. Numerical experiments support the versatility and computational tractability of this approach.
Keyword:
Approximation algorithms
Heuristic algorithms
Probabilistic logic
Dynamic programming
Function approximation
Convergence
Complexity theory
Continuous state-space Markov decision processes (MDPs)
dynamic programming (DP)
reinforcement learning (RL)
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7
论文数:
1.3W
被引数:
6.7W
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Simulation-based optimization of Markov decision processes: An empirical process theory approach
AUTOMATICA
IF5.9

