arrow
返回

Bootstrap model selection

delete1996-06-01
delete205
PRE
AI
J
Jun Shao *
DOI:10.2307/2291661delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In a regression problem, typically there are p explanatory variables possibly related to a response variable, and we wish to select a subset of the p explanatory variables to fit a model between these variables and the response. A bootstrap variable/model selection procedure is to select the subset of variables by minimizing bootstrap estimates of the prediction error, where the bootstrap estimates are constructed based on a data set of size n. Although the bootstrap estimates have good properties, this bootstrap selection procedure is inconsistent in the sense that the probability of selecting the optimal subset of variables does not converge to 1 as n --> infinity. This inconsistency can be rectified by modifying the sampling method used in drawing bootstrap observations. For bootstrapping pairs (response, explanatory variable), it is found that instead of drawing n bootstrap observations (a customary bootstrap sampling plan), much less bootstrap observations should be sampled: The bootstrap selection procedure becomes consistent if we draw m bootstrap observations with m --> infinity and m/n --> 0. For bootstrapping residuals, we modify the bootstrap sampling procedure by increasing the variability among the bootstrap observations. The consistency of the modified bootstrap selection procedures is established in various situations, including linear models, nonlinear models, generalized linear models, and autoregressive time series. The choice of the bootstrap sample size m and some computational issues are also discussed. Some empirical results are presented.
Keyword:
autoregressive time series
bootstrap sample size
generalized linear model
nonlinear regression
prediction error

期刊

J
Journal of the American Statistical Association
IF:
3
论文数:
5.2K
被引数:
4.8W

机构

暂无机构信息
引用论文

引用论文

暂无论文信息