Return
Cross-validation with antithetic Gaussian randomization
S
S
S
DOI:10.1093/jrsssb/qkag073.png)
Abstract
En 中文
We introduce a new cross-validation (CV) method based on an equicorrelated Gaussian randomization scheme. Our method is well-suited for problems where sample splitting is infeasible, either because the data violate the assumption of independent and identically distributed samples, or because there are insufficient samples to form representative train-test data pairs. In such problems, our method provides a simple, principled, and computationally efficient approach to estimating prediction error, often outperforming standard CV while requiring only a small number of repetitions. Drawing inspiration from recent splitting techniques like data fission and data thinning, our method constructs train-test data pairs using Gaussian randomization. Our main contribution is the introduction of an antithetic Gaussian randomization scheme, involving a carefully designed correlation structure among the randomization variables. We show theoretically that this antithetic construction can eliminate the bias of CV for a broad class of smooth prediction functions, without inflating variance. Through simulations across a range of data types and loss functions, we demonstrate that our estimator outperforms existing methods for prediction error estimation.
Keywords:
antithetic sampling
cross-validation
data fission
data splitting
model selection
variance reduction
Journal
J
IF:
3.6
Papers:
1.5K
Citations:
3.2W
