arrow
Return

Off-Policy Confidence Interval Estimation with Confounded Markov Decision Process

delete2022-10-05
delete5
delete
OA
AI
C
Chengchun Shi *
Z
Zhu, Jin
Y
Ye, Shen
L
Luo, Shikai
Z
Zhu, Hongtu
R
Rui Song
DOI:10.1080/01621459.2022.2110878delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article is concerned with constructing a confidence interval for a target policy's value offline based on a pre-collected observational data in infinite horizon settings. Most of the existing works assume no unmeasured variables exist that confound the observed actions. This assumption, however, is likely to be violated in real applications such as healthcare and technological industries. In this article, we show that with some auxiliary variables that mediate the effect of actions on the system dynamics, the target policy's value is identifiable in a confounded Markov decision process. Based on this result, we develop an efficient off policy value estimator that is robust to potential model misspecification and provide rigorous uncertainty quantification. Our method is justified by theoretical results, simulated and real datasets obtained from ridesharing companies. A Python implementation of the proposed procedure is available at https://github.com/Mamba413/cope. Supplementary materials for this article are available online.
Keywords:
Infinite horizons
Off-policy evaluation
Reinforcement learning
Ridesourcing platforms
Statistical inference
Unmeasured confounders

Journal

J
Journal of the American Statistical Association
IF:
3
Papers:
5.1K
Citations:
4.8W

Organization

L
London School Economics and Political Science
Scholars:
3.8K
Papers: 3.2K
Citations: 40
U
university of north carolina
Scholars:
7.4W
Papers: 6.5W
Citations: 93
S
Sun Yat Sen University
Scholars:
9.9W
Papers: 7.2W
Citations: 95
U
university of london
Scholars:
21.5W
Papers: 19.7W
Citations: 305
N
North Carolina State University
Scholars:
2.6W
Papers: 2.3W
Citations: 3.7W
researcher View more organizations