arrow
Return

Adaptive welfare maximization

delete2026-02-01
delete1
PRE
AI
P
Ponomarev, Kirill *
S
Shi, Liquiang
DOI:10.1007/s42973-026-00242-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We consider the problem of learning optimal treatment policies from observational data. We propose an algorithm that combines doubly robust welfare estimation, to accommodate rich covariates and unknown propensity scores, and sample splitting, to adaptively select policy complexity. We show that the resulting treatment rule achieves the minimax-optimal rate of convergence in expected regret while selecting a suitable policy complexity with nearly oracle performance. Our analysis avoids unnecessarily restrictive assumptions commonly imposed on the data-generating process or on first-stage nonparametric estimators and yields a sharp characterization of the relevant universal constants. The practical performance of the proposed method is demonstrated in a simulation study.
Keywords:
Double robustness
Empirical welfare maximization
Minimax regret
Model selection

Journal

J
Japanese Economic Review
IF:
0.5
Papers:
36
Citations:
0

Organization

U
university of chicago
Scholars:
4.4W
Papers: 3.7W
Citations: 80