arrow
Return

Gradient-bounded dynamic programming for submodular and concave extensible value functions with probabilistic performance guarantees

delete2022-01-01
delete2
delete
OA
AI
D
Denis Lebedev *
P
Paul J. Goulart
K
Kostas Margellos
DOI:10.1016/j.automatica.2021.109897delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We consider stochastic dynamic programming problems with high-dimensional, discrete state-spaces and finite, discrete-time horizons that prohibit direct computation of the value function from a given Bellman equation for all states and time steps due to the curse of dimensionality . For the case where the value function of the dynamic program is concave extensible and submodular in its state-space, we present a new algorithm that computes deterministic upper and stochastic lower bounds of the value function in the realm of dual dynamic programming. We show that the proposed algorithm terminates after a finite number of iterations. Furthermore, we derive probabilistic guarantees on the value accumulated under the associated policy for a single realisation of the dynamic program and for the expectation of this value. Finally, we demonstrate the efficacy of our approach on a high-dimensional numerical example from delivery slot pricing in attended home delivery. (c) 2021 Published by Elsevier Ltd.
Keywords:
Dual dynamic programming
Function approximation
Real-time operations in transportation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Automatica cover
Automatica
IF:
5.9
Papers:
1.2W
Citations:
5.2W

Organization

U
university of oxford
Scholars:
9.7W
Papers: 8.6W
Citations: 137